Human-led.
AI-amplified.
AI lowers the cost of writing code; it does not lower the cost of being right about the code. Our practice puts senior judgment at the center of every AI-assisted decision and treats every model output as something we have to defend, not just deliver.
How we work
with AI.
AI amplifies expertise. It does not replace it.
Senior judgment, domain knowledge, and client relationships remain the foundation. AI extends the reach of the practitioner; it does not stand in for the practitioner. Speed without judgment is not a service we offer.
Human review on every AI-assisted change.
Every AI-touched production change carries a sign-off from a senior practitioner who can defend it at audit. If a decision the model makes cannot be traced to inputs and policy, it is not a deployable decision.
Eval first, not eval later.
Evaluation harnesses are built alongside the model, not after. The team operating the system runs the evals; we hand them off as a deliverable, we do not park them as an internal tool.
No black-box defaults.
Where we use AI, the trail from input to output is documented and traceable. Where it cannot be, we choose a different tool. Auditability is non-negotiable on regulated work.
Where AI shows up
in our work.
Code generation, with senior-practitioner review as the gate. A meaningful portion of new code in a typical engagement is AI-assisted at first draft; one hundred percent of it is human-reviewed before merge.
Documentation, summarization, and structured extraction — places where the cost of being wrong is recoverable and the speed-up is large. Drafts get reviewed; we do not ship un-reviewed model output.
Evaluation harness construction, test scaffolding, and refactoring — repetitive work where the model strengths align with the deliverable. Each output sits in a pull request that a person who can defend it has read.
Where AI does not
show up.
Production decisions a human cannot defend in plain language at audit. If the answer to "why did the system do that?" is "the model said so," the answer is not deployable.
Architectural decisions on systems we will hand off. Architecture carries institutional weight that lives in senior practitioners, not in training data. We use AI as a sounding board; the call belongs to a named person.
Anything regulatory or contractual that requires a named accountable human. We do not let AI sign off on what a person needs to be on the hook for.
Three surfaces
that anchor trust.
A human-review surface.
Every AI-assisted output sits in a queue a senior practitioner reads before it ships. The review is the work, not an optional checkbox. Throughput is capped by review capacity, by design.
An evaluation harness.
Built alongside the model, owned by the engagement, handed off to the client team. Drift detection, regression checks, and the canonical eval set are documented artifacts, not tribal knowledge.
A traceability log.
Every model decision in a production system can be traced to its inputs, the policy it was acting under, and the version of the model that produced it. The auditor gets a real answer, not a reconstruction.
The bottleneck
has moved.
Industry research is consistent: AI-augmented teams ship roughly twice as much code, and spend nearly twice as long reviewing it. Judgment, context, and communication are now the scarce resources. We staff for where the bottleneck is now.
What buyers ask
about our AI use
Are you using AI in our codebase right now?
If your engagement is underway, yes. AI-assisted code generation is part of how our senior practitioners work; every output is human-reviewed before merge. The presence of AI in the workflow does not change who is accountable for the code: a named senior practitioner who can defend the decision at audit.
Can we restrict where AI is used in our engagement?
Yes. Some clients have explicit policies (data residency, model provider restrictions, classes of code where AI is prohibited). We document the restrictions in the engagement brief and operate inside them. If a restriction makes the engagement infeasible, we flag it before kickoff, not afterwards.
Do you store our code or data with the AI provider?
We default to providers and configurations that do not retain prompt or completion data for training. Where your environment requires stricter controls (on-premise inference, dedicated tenants, BYO model), we adapt. The data-handling shape is part of the engagement brief, signed off at kickoff.
What happens when the model makes a mistake?
It usually never reaches production because the human-review surface catches it. When something does slip through, the traceability log lets us reconstruct what the model saw, what policy it was acting under, and which review missed it. That goes into the eval set so the same mistake does not ship twice.
How is this different from your AI Solution Building service?
This page describes how AI shows up in every engagement we run. AI Solution Building is a service where the AI capability is itself the deliverable. Same principles apply; the AI service line is what you hire us for when you want to build an AI-powered system. See the services page for the service-level detail.
Want the
longer read?
The bottleneck has moved is the firm's full essay on AI-augmented delivery. The AI governance checklist is a working tool for regulated environments. Ask us for either one. Or skip ahead and send us a problem statement.
letscreate@nuarch.com