All field notesWorkflow Ownership

Your AI Moat May Be the People Who Can Grade the Answer

Before an SME debates building or buying AI, turn ten difficult past cases into a grading set that domain experts can defend.

Access to another AI model is not a moat.

The harder asset to copy may be the people who can look at an answer, explain why it is wrong, and define what “good enough” means for the business.

That matters because many companies begin an AI project with the wrong question:

Should we build a model or buy one?

For most SME leaders, that question arrives too early. Before choosing a model, vendor, or platform, the business needs a way to judge whether any of them can do the work.

A recent Thomson Reuters example makes the distinction useful. The Decoder reports that the company spent about US$40 million over more than two years building a legal model on Qwen, its proprietary content, and input from hundreds of domain experts. The same report says its advantage was strongest when the system could use exclusive company content. In the cited web-only factual-accuracy comparison, the result was weaker.[2]

That does not prove every company should build its own model.

It points to a narrower lesson: the model was only one part of the asset. The company also had specialised evidence and people capable of judging difficult legal work.

For an SME, the practical starting point is much smaller.

Do not begin by training a model. Begin by building a grading set.

Turn ten difficult cases into an AI test

Choose ten completed cases from a recurring business workflow.

Not ten easy cases. Select the ones that required judgment:

  • an invoice with a disputed line item;
  • a sales opportunity with incomplete CRM history;
  • a customer complaint that crossed service and billing teams;
  • a contract clause that needed escalation;
  • a stock exception where the system record and physical count disagreed.

Remove private information that the test does not need. Preserve the context that made each case difficult. Then write down what a competent person should notice, decide, explain, and escalate.

This becomes the test set for a vendor demonstration, an internal prototype, or a digital coworker.

A polished demo asks the AI to perform a task the seller already knows it can complete. A grading set asks whether the system can handle work that has already caused friction inside your business.

That is a much better buying test.

Define observable quality

“Looks good” is not a scoring method.

Ask a domain expert to convert quality into observable criteria. For each case, score the answer on four dimensions.

1. Correctness

Did the system identify the facts, rules, calculations, and constraints that affect the decision?

A fluent explanation can still be wrong. The score should depend on the answer matching the evidence and the business rule, not on how confident it sounds.

2. Completeness

Did it cover the required parts of the job?

A finance agent might identify an invoice mismatch but fail to check the approval status. A CRM agent might draft a sensible follow-up while ignoring the customer's unresolved support issue. Partial work can look useful while creating downstream risk.

3. Acceptable uncertainty

Did the system recognise when the evidence was incomplete or conflicting?

The best answer is not always a decision. Sometimes it is a clear statement that the case cannot be completed safely without another document, field, or human judgment.

4. Escalation

Did it send the exception to the right person with enough context to act?

“Human review required” is not enough. A useful escalation names the issue, shows the supporting evidence, explains what the system attempted, and states the decision that remains open.

These four criteria turn domain knowledge into system requirements. The expert is no longer only a user of AI. The expert is helping to architect how the AI should behave.

Use two graders, not one

Have at least two qualified people score the same outputs independently.

Their disagreements are valuable.

If one finance lead accepts an answer and another rejects it, the immediate problem may not be the AI. The business may have an unwritten rule, inconsistent practice, or a quality standard that nobody has made explicit.

Document the disagreement. Ask what evidence changed the judgment. Decide whether the rule should become part of the rubric, remain a human decision, or trigger escalation.

This produces three useful assets:

  1. a set of representative cases;
  2. a scoring rubric the business can defend;
  3. a record of where expert judgment is still required.

A vendor cannot create those assets from a generic product demonstration. They come from the way your company actually works.

Proprietary data is not enough

Companies often hear that their data is their moat.

That is incomplete.

A folder full of historical documents can contain duplicates, stale rules, weak outcomes, and inconsistent decisions. Feeding it into an AI system does not automatically create an advantage.

The useful asset is curated evidence connected to a quality decision:

  • this case was accepted, and here is why;
  • this answer failed, and here is the missing evidence;
  • these two experts disagreed, and this is the escalation rule;
  • this uncertainty is acceptable, but that uncertainty stops execution.

That combination is harder to copy than raw data. It reflects how the business distinguishes a plausible answer from a dependable one.

Use the grading set before you buy

Give the same ten cases to every shortlisted vendor or internal system.

Keep the rubric stable. Record the model and configuration used. Compare accepted outputs, critical failures, unsupported claims, escalation quality, time, and cost.

Do not let a vendor replace your difficult cases with a cleaner benchmark. Do not score only the average. One high-risk failure may matter more than nine good answers.

The result will not tell you whether to build or buy in every situation. It will give you evidence for the next decision:

  • Is an existing product already good enough?
  • Does the system need retrieval from company records?
  • Which errors can rules catch?
  • Where must a person approve the outcome?
  • Is customisation worth the added cost and maintenance?

That is a more useful conversation than arguing about model brands.

Start with judgment, then add technology

The AI market will keep making capability cheaper and easier to access.

Judgment will not become cheap at the same rate.

A domain expert who can define correctness, expose missing context, resolve disagreements, and design escalation is doing more than reviewing AI output. That person is turning operational knowledge into an architecture the system can follow and the business can audit.

Before your next AI procurement meeting, pick ten hard cases.

If your team cannot agree on how to grade them, another model will not solve the problem.

Sources

[2] https://the-decoder.com/thomson-reuters-bets-40m-on-owning-its-ai-instead-of-renting-from-openai-or-anthropic — Thomson Reuters bets $40M on owning its AI

Continue the work

Turn a capable model into dependable execution.

Nexius Labs helps SMEs design the context, tools, permissions, approval gates, and evidence trails around useful Digital Coworkers.