Insurance

Claims intake that does not stall at the inbox

Claims intake AI automation for a mid-market insurance carrier: classification, extraction, and a human-in-the-loop queue. Wrong routing delayed a payout, so the model only routes what it can defend.

Result

62% less manual review

What we left alone

We did not touch liability decisions. Those still belong to a licensed human. The system earns more of the sorting job over time; it does not take the judgment.

The situation

This was claims intake AI automation for a mid-market insurance carrier—not a chatbot, not a lab pilot. Incoming claims arrived by email, PDF, and portal. Adjusters spent mornings classifying, extracting, and routing instead of deciding. Leadership had already bought an “AI claims” demo that classified five happy-path PDFs and then sat unused. The function that hurt was intake: pull the right fields, send the obvious cases forward, and put the rest in a human-in-the-loop queue with the evidence attached. Wrong routing delayed a payout. The team would not accept a black box.

The real constraint

The work was not “understand a claim.” The work was production claims intake: extract, classify, route what the model can defend, and park the rest in front of a licensed person with the evidence attached. Wrong routing delayed a payout. The team would not accept a black box.

Approach

What we did instead of the obvious demo

01

We threw out the chatbot

The previous vendor had built a chat window. Adjusters do not chat with intake. They need a queue. We designed a desk: extract, classify, confidence, next action.

02

The eval set was the arguments

We did not test on clean samples. We built the set from claims reviewers had already disagreed about. If the system could not explain those, it was not ready.

03

Low confidence is a feature

Below the threshold, nothing ships. A person sees the document, the extracted fields, and the reason the model hesitated. That queue is the product as much as the model is.

04

Owned by intake, not by “the AI team”

An intake lead owns the weekly review. They can raise the threshold on a Friday without a deploy. That is what made it last after we left.

What we refused

  • End-to-end “auto-adjudication” in the first version
  • Training on easy claims that no one argued about
  • A model that wrote customer-facing language

What transferred

  • The queue, the thresholds, and the weekly review ritual
  • An eval set the carrier still adds to
  • A runbook for document types the model has never seen

The outcome

After twelve weeks in production, 62% of incoming claims left the inbox without a human sort. Average time-to-desk on the remainder dropped because the evidence was already attached. Error on routed work stayed inside the band the intake lead had signed. The claims intake AI automation still runs without us—that is the difference between a production system and a pilot.

Tell us the job that already hurts.

We’ll tell you honestly if we can help and what it would take.