← Notes from the work

October 2026 / Qazeem Oladejo

In regulated AI, limits are a feature

Why source verification, visible limits and human review mattered to legal teams considering an AI assistant.

I led product on an enterprise LLM assistant for legal teams. It sits inside Microsoft Word and Outlook, and it helps with drafting, compliance review, fact-verification and regulatory-filing analysis. It was built for a client of the studio I worked at, by a team of three engineers.

A central product decision was where the assistant should stop. We defined those limits, made them visible and accepted the delivery costs they introduced.

Buyers were not asking for more

I ran discovery workshops and interviews with enterprise users, buyers, and compliance and legal stakeholders. I wanted to understand their buying criteria and, more usefully, what stopped them adopting. Three objections came up.

  1. It might invent a fact or a citation. If it does, the lawyer carries the liability, not the software.
  2. Client documents might leave the firm, or be used to train a model.
  3. There would be no record of what the tool did, so its work could not be defended later.

The objections concerned accuracy, confidentiality and accountability. Adding drafting or summarisation features would leave those concerns unresolved.

In a regulated field, the buyer's first question is what the product will never do.

Designing the refusals

So we designed the limits as carefully as the features. The assistant has three refusals.

  • It will not give a legal conclusion with no source to support it.
  • It will not answer from outside the documents in the user's workspace.
  • It will not send or file anything externally without a person approving it.

Each refusal maps to an objection. The first answers the fear of invented citations: outputs are grounded in a retrievable source, and the user can verify that source for themselves. The second answers the fear about client documents: each workspace's documents stay inside that workspace, with admin controls over how the assistant is used. The third answers the fear about accountability, together with human-in-the-loop review and a full audit trail of what the system did and why.

Those six parts, citations, source verification, workspace data boundaries, human review, the audit trail and admin controls, became what we called the trust layer.

A limit has to be visible

A refusal that happens silently is a bug from the user's point of view. So I also defined how the product's limits were communicated. When the assistant declines, it says so, and the user can see why.

This matters more than it sounds. A tool that declines openly is one a lawyer can plan around: they know when to rely on it and when to do the work themselves. A tool that guesses gives no such signal. After the first wrong answer, they stop trusting any of its answers, and quietly stop opening it.

What it cost

These decisions had practical costs in speed and engineering capacity.

Mandatory human review added time to the workflow. We kept it because users needed control over work sent or filed in their name.

With three engineers, building the audit trail, workspace boundaries and admin controls reduced capacity for other features. I prioritised these controls because they addressed the adoption barriers buyers had described.

Agree the rules before building

Three groups had to agree before any of it was built: the users, the compliance stakeholders and engineering. I facilitated that agreement on data handling, human review and the audit trail.

It would have been faster to build first and present it for approval. It would also have risked a rebuild after a compliance review. Agreeing early was the cheaper path.

I learned the legal workflows and documented the requirements agreed with stakeholders. Engineering then had a shared reference for decisions and questions.

Quality as a release gate

Trust is not established once. A model's behaviour changes between versions, so a prompt that was safe last month may not be safe this month. We treated output quality as a gate on every release, alongside QA, security, privacy and risk reviews.

  • A fixed set of test prompts ran on every release.
  • Outputs were scored for accuracy, and for whether each claim traced to a source.
  • A lawyer or legal stakeholder reviewed a sample before release.
  • Quality was tracked over time.

Each full release was also preceded by alpha and beta programmes with real users. I presented demos to client stakeholders, and spent a good part of them explaining model behaviour and limits to non-technical audiences.

What I would tell another product team

  • Start discovery with the reasons to say no. They are the roadmap.
  • Write down what the product refuses to do before you write down what it does.
  • Make every limit visible to the user at the moment it applies.
  • Expect trust work to cost speed and features, and decide in advance that you will pay.
  • Test the model on every release as if it were new, because in effect it is.

Legal teams adopted the assistant. My strongest lesson was that clear limits and review controls can be worth the time they add, especially when users remain accountable for the output.