What agentic AI is still taken for: three lessons from Stockholm
From a keynote to Nordic financial institutions: why agentic AI is not a black box, why kill switches miss the point and how accuracy can rise.
At a recent event in Stockholm I gave a keynote to Nordic financial institutions. The two frameworks behind many production AI agents, Pydantic AI and LangChain, reached version 1 only in autumn 2025. At Vstorm we had been shipping agentic systems to production before that starting line.
On stage I gave an engineer's view of what production-grade agentic AI actually is, and in the conversations afterwards I learned how far that view sits from the one decision-makers hold. There is less wishful thinking than a year ago, but the picture of what AI can do is still mostly off. Here is the gap between what agentic AI is and what businesses assume it to be.
AI is seen as an abstract model, a black box at work #
Three years into a world with capable LLMs, most decision-makers I spoke to are still unsure what goes into agentic AI. They picture a non-deterministic language model in the middle that pulls all the strings.
What they miss is that every agentic AI system combines a stochastic core with deterministic scaffolding built around it. I usually open my keynotes with a slide that separates the two: the LLM core and the deterministic shell around it.
The scaffolding intercepts, structures and acts on everything that reaches the model and everything the user gets back, and engineers have full control over it. Users are not at the mercy of whatever the model returns, large or small. It is the engineer's job to validate the outcomes and protect against edge cases.
Talk about kill switches steals the spotlight #
Because it is still unclear what an AI system is, people tend to see these systems as uncontrollable and potentially dangerous. And what we fear and cannot control, we want a way to shut down.
From an engineering perspective, this view is deeply flawed. It is as if toddlers who can barely walk were discussing the risks of high-speed skateboarding.
Most production-grade systems are not autonomous in 2026. They support people by taking the mundane work off their shoulders. You do not need a kill switch for a system that processes claim documents ahead of a human decision, or for one that suggests how to allocate investment funds.
These systems can make mistakes, and it is the engineers' job to guard against them with the scaffolding they build around the LLM. In our projects, what ships to production works in over 90% of cases. If a system removes manual work in 9 situations out of 10, that is a reason to celebrate, not to shut it down, and a good base for improving accuracy before anyone talks about fully autonomous processes.
AI cannot possibly aim at 100% accuracy #
Reasonable thinking leads to the wrong conclusion here. If AI is probabilistic at its core, the argument goes, it can never be fully reliable, because that is its nature.
On our most ambitious projects at Vstorm we already see that this is not necessarily the case. The view starts to shift once the LLM is treated as one tool in the toolbox rather than the ultimate answer.
Every solution we build explores a range of possible situations. If that range is mapped out and deeply understood, a classical deterministic system can work through it with a rulebook: if this happens, do that. Where that understanding is missing, an LLM does the job better, because it can move through an undefined space without a rulebook and spot patterns people may not have seen.
What we have found on projects is that the longer an LLM works through such an undefined space, the better the client organisation understands it, by watching and learning from the AI at work. That creates a chance to write or update the rulebook, and for some parts of the job the solution can then fall back to deterministic systems.
Whether this works for most projects is still an open question, but some projects in the Vstorm portfolio already show the trend. We have found that 100% accuracy can be within reach, not through some superintelligent AI, but through a well-engineered solution that uses both deterministic and non-deterministic components for the problem at hand.
Conclusions #
The engineer's perspective differs from conclusions reached by reasoning alone. One thing everyone at the conference agreed on is that working on real projects is how you learn to talk about AI put to work.
It is also the hardest perspective to get, because you have to take on the risks of adoption work, and hands-on work with AI does not give you the clean answers that reasoning does. In the end, though, the conclusions earned on real adoption projects are what can save others the time and money they would spend learning the same things themselves.
In the keynote I also presented conclusions from projects Vstorm delivered to financial institutions. You can read them in the follow-up post, Our findings in agentic automation from the Nordics Summit. If you work at a financial institution and want to adopt agentic AI the right way, you can book a consultation with Vstorm.


