Case Study

Agentic AI claim processing: from hours of manual review to minutes of verified analysis

A US healthcare insurer processing complex, multi-document accident claims asked Vstorm to cut administrative burden and error in review. The system combines two LLMs, algorithmic validation, LlamaParse, and a live benefits API — hours to minutes, with a human still making the decision.

  • Healthcare / Insurance
The outcome

Three hours of reading, down to an eight-minute verified pack

Accident claims arrive as policies, medical notes, accident narratives, scans, and sometimes a police report. The reviewer had to align dates, coverage, exclusions, and what the benefits database already paid. The agent does that assembly. It does not approve or deny.

3 hours → 8 minutes is this insurer's published processing time, not a STCC clinical score and not a Medicare pre-appointment hour. Two models (GPT and Gemini) plus an algorithm: agree, and confidence is high; diverge, and a human looks.

To process one claim by hand

3 hrs

Dates, exclusions, benefit limits, handwriting, and overlapping policies.

2

LLMs run in parallel

GPT and Gemini. A third algorithmic layer validates both.

4

Data sources on every claim

Digital policies, scanned forms, images or maps, and the benefits API.

About the client

A leading US healthcare insurance company, not named here. It offers multiple plan types for different financial circumstances and coverage needs. A large share of claims are accident-related. The book also covers maternity, non-accident medical treatment, and other benefit categories.

Vstorm's impact

Vstorm's impact, the TL;DR

  • Claim processing cut from 3 hours to 8 minutes, with a specialist still deciding
  • GPT and Gemini in parallel, plus algorithmic validation — a triple check
  • LlamaParse on digital files and scans; images and maps when the file has them
  • Live API into the benefits database so limits already used are not guessed
  • Structured summary: dates, exclusions, benefit status, and flags for review

The challenge

Dates, handwriting, exclusions, and a second review on appeal

The 2025 State of Claims report found 41% of respondents saying at least one in ten insurance claims is denied. Of those denials, 26% are blamed on inaccurate or incomplete intake data — errors that could have been caught earlier. This insurer also carries extra load when a patient asks for reevaluation: a second cycle, more paper, more strain.

The target is a claim that is clearly approved or clearly denied on complete information, with little ambiguity left for an appeal. Accident files make that hard. The pack can include the policy, medical procedure notes, an accident description, a police report, scans, and sometimes a map showing why emergency transport was needed.

From that pack the reviewer must check whether cover was live at the time of the accident, whether the procedures sit inside limits, and whether exclusions apply — DUI, or an X-ray without a qualified request, among others. Dates compete: accident, birth, policy start and end, email timestamps, visit dates. Overlapping policies (private, employer, veteran) make the relevant date a relationship, not a string. Handwriting confuses a 7 with a 1. Paper tears and creases. Rule engines alone cannot do the contextual exclusion work.

How we delivered

TriStorm on a file the specialist still owns

The agent assembles a verified pack. A human still says pay or not. The same workflow is being extended to maternity and hospital claims.

Map the accident file

Workshops locked the four sources — digital policy, scan, extra modalities, benefits API — and the failure modes: date fog, handwriting, policy-specific exclusions, overlapping cover.

  • Document inventory
  • Date and exclusion brief
  • Benefits-API gap

Proof of Value on dual-LLM review

LlamaParse structures the file. GPT and Gemini read independently. An algorithm validates both. Agreement is high confidence; divergence is a flag, not an auto-merge. Fine-tuned prompts extract exclusions and compare them to the case.

  • LlamaParse path
  • GPT + Gemini parallel pass
  • Algorithmic third check

Azure, CosmosDB, a human at the end

Microsoft Azure, CosmosDB, LlamaParse, GPT, Gemini, and the existing benefits API. The output is a summary for a specialist. The stack is being extended to more policy types with little re-architecture.

  • Azure + CosmosDB
  • Reviewer-ready summary
  • Extra policy-type track

How it works

Four sources, two models, then a person

Each claim runs from raw intake to a structured summary. Four sources: digital policies, scanned physical forms, supporting images or maps, and a live API to the benefits database — so monthly, quarterly, or yearly procedure limits are facts, not a guess from the PDF.

Undisclosed healthcare insurer — claim processing workflow
Claim workflow as built — Vstorm × undisclosed US healthcare insurer
Claim intake Digital policies, scans, images, live benefits API
Document parsing LlamaParse structures the file
Dual-LLM analysis GPT and Gemini independently — agree or flag
Algorithmic validation Third QA layer on the model outputs
Benefits database Live API: what was already paid, what limits remain
Human-reviewed summary Dates, exclusions, benefit status — specialist decides

Vstorm × undisclosed US healthcare insurer — agentic claim pipeline

The system does not approve or deny claims. It prepares the specialist to do so with complete, verified information.

Results

What changed on the clock

Claim processing time
~22× faster
Manual review
3 hrs
With the claim agent
8 min

The same workflow is being extended to maternity care and hospital treatment. The stack stays Azure, CosmosDB, LlamaParse, GPT, Gemini, and the client's benefits API.

Work with us

Ready to see how agentic AI transforms claim-review workflows?

Meet directly with our founders and PhD AI engineers. We will walk through real implementations from 30+ agentic projects and the practical steps to integrate them into your workflows.