From expensive chat buddy to agentic systems

5 Levers for Strategic AI Deployment of agentic systems.

7 min read Translated from German
Sewing cotton
Image by Ian Taylor / Unsplash

5 Levers for Strategic AI Deployment

Many companies now use Large Language Models (LLMs) quite intensively in their daily work. Often, however, the introduction of the shiney new AI technology was uncoordinated and ad-hoc, driven by motivated individuals or pilot teams. The result: companies now face high costs while the actual value remains unclear or hard to measure.

The reason usually lies in the approach: results are often generated step-by-step through a discursive chat process. This process is a one-off effort, and once the document is finished, its long-term reusability is lost. While this may shorten creation time of documents or results, it is inefficient over time and won't scale.

The strategic approach is: Automate recurring tasks and queries once, whether through prompts, skills, or agents.

1. Improve the Prompt, Not Just the Output

This is a no-brainer for frequent routine tasks, but it is equally useful for occasional work: Do not work toward the final result directly with the LLM chatbot; instead, work on the prompt, the necessary context, and open questions. Then create the output.

If the output is not as desired, you don't just work with the output adding something here and there. Instead, add the missing requirements directly into the original prompt. Step by step, this builds a workflow that can be re-used many times and runs automatically at the push of a button next time.

Use Prompting Patterns and good practices as communicated by the model operators

Around 70% of users (I might have made this data up) prompt as if they were doing a standard web search or talk to a friend at the bar. While this works okay in many cases, it leaves a huge amount of potential on the table. A solid handful of prompting patterns raises your prompts to a completely different level.

Beware of Over-Specification

Continuously inflating prompts for every edge case makes them fragile and vulnerable to model updates (the prompt drift). The technology changes fast and so do the standards. Remember CustomGPTs? Exactly that.

Prompts for language models are not deterministic software code that creates the same output each and every time, and they do not behave like it. Overly complex prompts and layers of context become expensive fast, and the model will start weighing context rules and contradicting information against each other (or, very likely, ignoring them entirely). The big model operators like Anthopic or OpenAI frequently release solid documentation on the current best practices for their models. Reading through them is besides all marketing usually worth your time.

This first lever also applies to agentic systems.

2. Automate Real Workflows and Use-Cases

An old saying of the good old days of digitalization is “A digitalized bad process is still a bad process.” This applies without exception to the current wave of digitalization using the AI of your choice. Therefore, always clarify and define the use case and the actual process first, then automate.

However, there is a caveat when going fully autonomous: Agent maintenance must be factored in. Creating an agent is only 20% of the work, the remaining 80% is maintenance, monitoring, incident handling, and ongoing upkeep. This mistake is still common in software development, let's be smarter now.

3. Use Tokens Economically

The era of heavily subsidized AI tokens being thrown into the market is slowly but surely coming to an end. The clever use of tokens is thus becoming a core capability within your company.

Always start with the smallest and cheapest model. Only scale up the model size when needed, such as when the output quality falls short. But do keep in mind: Small, budget-friendly models break quickly when faced with complex prompting patterns, multi-step logic, reasoning, self-reflection, or strict format adherence.

Therefore, you need to know your models. Learn their pros and cons (e.g., via an OpenRouter budget for experiments with different models). Not every task requires a massive flagship model — many sub-steps can be handled by decision models (like Jev and similar) or smaller, specific models. This preserves budget, which pays off significantly at scale.

My rule of thumb for economic token usage:

Use small models for structured sub-tasks (classification, extraction, summarization) and large or smart models for orchestration, complex reasoning, and quality checks. Use low-cost decision models wherever possible (quality checks within bandwidths and treshholds, simple decision points).

4. The Coordination Trap: What Should Actually Be Automated?

Not everything that is technically possible needs to run in auto-mode. Often, the largest lever lies in using AI for recurring coordination tasks: routines, consistent processes, meeting ops or structured preparation of the next presentation extravaganza. This allows human attention to stay focused where judgment, responsibility, and complex alignment are required.

Automation in the AI era is, well, classic product management:

  • Analysis: Model your domain and processes; identify interfaces and multi-step tasks.
  • Prioritization: Evaluate use cases by impact, risk, and effort.
  • Specification: Define the objective, ideal workflow, and expected artifact.
  • Execution in agile ways-of-working: Start small with clearly scoped workflows and defined risks, then iterate from there.

5. Set Guardrails for Agents

The more autonomous an agent or workflow is allowed to operate, the more critical clear guardrails become:

  • Limited Permissions & Stop Rules. Build in clear authorizations, explicit abort criteria, and human checkpoints (Human-in-the-Loop). Consider error tolerances: If a human has to verify every single agent output for facts and hallucinations, the ROI drops significantly. True automation requires defining clear error tolerances.

  • Solve the Context Problem. Keep all required documents ready and structurally well-prepared. Explicitly and rigidly exclude irrelevant information.

  • Avoid Context Window Fatigue. Supplying excessive context consumes massive amounts of tokens, driving up costs. Additionally, it often degrades model accuracy (the Lost in the Middle phenomenon).

  • Keep Prompt & Agent Complexity Manageable. Split the process into smaller, modular sub-steps rather than building huge chains. Recent developments are shifting in this direction (prompting loops and coordinating agents).

  • Embedd Transparency. Make the LLM display facts, conclusions, and assumptions in its output so that processing steps remain traceable and quality checks part of the process.

  • Establish Quality Assurance. Set up systematic tests (evals), just as it's good practice in software development. Use checklists with real criteria, examples, and failure scenarios.

  • Treat Agents as Software. An agent or automated prompt is like a piece of software — it needs an dedicated owner, proper documentation, and regular maintenance. Of course, the maintenance can be agentic, too.

  • Account for Compliance and Data Security. Clarify upfront what can run in the cloud, what needs to run locally or isolated, and which data must be anonymized. Regulations become a huge part of the non-functional requirements.

Conclusion: The Transformation from Chatting to Systems

Moving from sporadic and uncoordinated AI usage to generating real business value requires a mindset shift: Away from one-off dialogues, toward reusable and scaleable systems. Anyone using AI merely as a better search bar or text generator will generate high costs with limited benefit. Worst case, it ends up as AI slop tennis.

Conversely, those who structure workflows cleanly, systematically optimize prompts, and use tokens consciously will achieve true scalability. AI takes over coordination and boring prep work while humans retain control and final decision-making.

Do not view AI implementation as a technical experiment, but as continuous product management. That is how a potential cost trap becomes a sustainable productivity gain.

These insights represent the key takeaways from the Product Prompting Sprint within the seminar "Future Skills in Digital Product Management".

Sources


AI disclaimer: This article was translated from German to English with the help of Gemini 3.6 Flash.