The clinical content team was producing one guideline per reviewer per day. We built AI agents that changed that to ten without replacing a single reviewer.
This company is a leading healthcare technology firm specializing in payment integrity, working with some of the largest health plans in the country. At the core of what they do is clinical content — policies, guidelines, and rules that determine whether a claim should be paid, denied, or reviewed. That content needs to be accurate, explainable, and current. The process of creating it was entirely manual.
A clinician or content specialist would take a large policy or clinical reference document, read through it, and identify language that could form the basis of a denial criterion. From there, a guideline would be drafted, reviewed, approved, and eventually coded into a rule that runs against live claim data.
At that pace, every week of delay in getting a policy document turned into an active rule was a week of revenue sitting on the table. Their clients are major health plans managing hundreds of millions of claims, were waiting on content that could not be produced fast enough.
One person. One document. One day. That was the pace the team was working at. The backlog of documents needing review was growing faster than anyone could work through it and every document waiting in the queue was a rule that wasn't yet running against claims.
To understand what we built, it helps to first understand what we were replacing. The content production process had five distinct stages, and a human being was responsible for every single one of them. The two middle steps of identifying denial language and drafting guidelines were the bottleneck.
Steps 2 and 3, highlighted above, required deep document reading, pattern recognition across large bodies of text, and the ability to consistently extract structured meaning from unstructured clinical language. These are exactly the things AI does well.
We built a set of AI agents that take a policy document as input and identify denial criteria from specific passages of text that match the patterns of language typically used to define clinical rules. Instead of a specialist spending a full day reading a document to find one guideline, the agent reads the entire document and presents a set of identified candidates for review.
The reviewer's job changed fundamentally. They were no longer reading documents. They were reviewing AI-generated policies, approving the accurate ones, correcting the ones that needed adjustment, rejecting the ones that did not apply. That is a very different cognitive task. And a much faster one.
The biggest shift was not that the AI was faster at reading documents. It was that the human's job changed from reading to reviewing. That is a fundamentally more valuable use of a clinical expert's time and it is why adoption happened without anyone being told to adopt.
We deployed the first version in 12 weeks using a hybrid pod of 4 engineers and AI coding agents. The approach was zero-trust: the system started in shadow mode, with the AI generating candidates but humans making every final call. As confidence scores improved, autonomy expanded gradually.
The 10x figure is not about the AI being faster at reading. It is about the AI being more thorough. A person reading a 200-page clinical document under time pressure will miss things. An agent does not get tired. It does not skim. It surfaces every passage that matches the pattern, consistently, every time.
There was no mandate, no rollout campaign, no required training. Adoption happened because the tool worked and because the people using it could feel the difference immediately. Reviewers who had spent years reading through dense policy documents found that they were now doing something different and better. Their output went up. Their frustration went down.
The biggest shift was not speed, it was the nature of the work. Reviewers went from reading to reviewing. That is a more valuable, more satisfying use of clinical expertise. When people feel like they are doing better work, adoption takes care of itself.
We started with the AI in shadow mode, generating candidates while humans made every final decision. This gave reviewers time to build confidence in the outputs before autonomy expanded. In a regulated clinical environment, that trust-building is not optional.
The 10x improvement in guidelines identified was not because the agent reads faster. It is because the agent does not miss things. A person reading under pressure skims. An agent surfaces every relevant passage, every time. That consistency is where the value lives.
We spent very little time on formal change management because we did not need to. When something genuinely makes your job easier from day one, people tell their colleagues. Sixty reviewers in three months without a single mandate.
If your team is spending time reading documents when they should be spending it on judgment, we should talk.
Start a conversation See our Implement methodology →