Clinical triage — Schmitt-Thompson
Guideline-executing triage lifting a raw LLM from 44% to 98% on the open 50-scenario benchmark — a domain where a wrong answer is a patient-safety event.
Your AI pilot works in the demo. We get it to production.
We audit the pilot you already have, fix what blocks production and hand it back to your team — whether it was vibe-coded, built by another vendor or locked into a platform.
Three-agent product advisor guiding customers through print-order configuration.
Read case studyText-to-workflow agents building validated node graphs inside the platform.
Read case studyHIPAA-compliant guideline-executing triage — 44% to 98% on the open 50-scenario benchmark.
Read case studySupply chain intelligence agents reducing manual coordination overhead.
Production AI workloads moved to new hardware for on-prem LLM deployment.
Multi-agent financial coach: pattern extraction, gamified habit-building and two coaching personas with opposite attitudes.
A pilot proves that an agent can do the task once. Production needs it to do the task on real data, at volume, at a cost you can carry, and in code your team can change. Those are engineering problems, and they can be measured and fixed without starting over.
Every figure here is a published client outcome tied to the evaluation set it was measured against — the standard we suggest you hold any vendor to.
Guideline-executing triage lifting a raw LLM from 44% to 98% on the open 50-scenario benchmark — a domain where a wrong answer is a patient-safety event.
The client aimed for 80% accuracy to be usable. The delivered multi-agent product advisor exceeded it in production, with an 11.76% order increase on day one of the Australian launch.
Engineers moved from hours of tedious setup to minutes, through multi-step validation rather than a single generation pass.
Four ways a pilot gets stuck, and the work that moves it forward.
A prototype built fast with AI coding tools, cleaned up for production: tests, error handling, secrets out of the code, logging, and a structure your team can maintain.
Accuracy, latency and token cost measured on your own cases, then fixed where they break: retrieval, prompts, tool design, model choice and caching.
An engineering review of an agent system before you scale it, buy it or invest in it: architecture, evaluation, security, data handling and what it will cost to run.
Agents moved from a closed or no-code platform to code and infrastructure you own, with the same behaviour checked against your existing cases before the switch.
An audit first, then a fix plan with a go / no-go decision, then the work itself — measured against the same baseline from start to finish.
We read the code, run it against your real cases and write down what blocks production.
Each blocker gets a fix, an owner and an estimate, including the option to rebuild a part instead of patching it.
We make the changes, prove them against the baseline and hand the system to your team with runbooks.
Published outcomes, each linked to the full case study.

“My wish was to come to at least an 80% success rate in the workflow results, and by the time we finished the project, the success rate is, I believe, over 95.4% — so it definitely exceeded expectations.”
95.4%
Success rate in workflow results
Schmitt-Thompson · Clinical triage
Nurse-triage guidance where a wrong recommendation is a patient-safety event — staged retrieval and validation rather than one model answering in a single pass.
44% → 98%
raw LLM vs guideline-executing accuracy on the open 50-scenario benchmark
Not adjectives — memberships, partnerships and open-source work anyone can check.
The first AI consultancy in the foundation — contributing to the standards for how production agents get built.
Official Pydantic implementation partner, working with the framework since its beta versions.
Our agentic AI libraries are used by developers worldwide; we contribute to the tooling we deliver on, rather than only consuming it.
Book a session. We look at what you have and tell you what blocks it and whether it is worth fixing.