SERVICE · 05

Audit & rewrite

We look inside the AI you already run, find where it fails or wastes money, and give you a clear order of what to fix.

You have an AI feature in production: rising costs, answers that don't convince, users who avoid it. With access to the system we measure what really happens, and reach one decision per component: keep, fix, rebuild or remove.

FROM EVIDENCE TO DECISION

Example extract: components and numbers are anonymised and illustrative. The same data follows you down the page.

audit · ai-support-stack

Example · anonymised data

  • —embedding-pipeline—

    accuracy 71%, below baseline · loss 2.3 on the test set

    wrong answers on ~3 questions out of 10 · effort: medium

    REBUILDNew embedding model and document re-indexing
  • —cost-attribution—

    $4.3k/mo unattributed · 60% of the AI budget

    no way to know which feature costs what · effort: low

    ADDCost tracking per call and per feature
  • —prompt-routing—

    23% of calls go to the premium model · ~$1.2k/mo recoverable

    cost without quality gain on simple cases · effort: low

    FIXRoute simple cases to a cheaper model
  • —retrieval-eval—

    recall@5 = 0.91 · above target

    works as intended · effort: —

    KEEPNo change; stays as the test reference
We measure every component on real traffic and on test cases.

WHAT YOU GET

Not “a big report”: four things you can act on, all in a written report plus a walkthrough session with your team.

  1. 01

    The evidence

    For every component: what we measured, how, and what came out.

    In the example: 71% accuracy and $4.3k/mo nobody can attribute.

  2. 02

    The priorities

    An order by impact and effort, so you know where to start.

    In the example: cost attribution first. Low effort, and it unlocks everything else.

  3. 03

    The rewrite plan

    What to keep, fix, rebuild or remove, in which sequence, with estimated effort.

    In the example: rebuild the embedding pipeline, keep retrieval as the reference.

  4. 04

    Costs and benefits

    Where an intervention saves money or improves quality, with the estimate and the assumptions behind it.

    In the example: ~$1.2k/mo from routing, with the quality risk made explicit.

HOW WE WORK

Indicative duration for a mid-sized system; we confirm it once we've seen the scope.

  1. 01days 1–2

    Setup & access

    Access to the systems, map of the components.

  2. 02days 3–8

    Analysis & evidence

    Tests, metrics, in-depth review of every component.

  3. 03days 9–10

    Report & walkthrough

    Written report and a session with your team.

WHO IT'S FOR

You have an AI feature in production, costs are rising or results don't convince, and you want to know what to rebuild before putting more money into it.

WHO IT'S NOT FOR

You want to build AI from scratch (→ Operational product builds). You want a theoretical opinion without access to the system: without data there's no audit.

  • Indicative duration: 10 working days
  • Price agreed before we start

Timelines and terms are indicative. We put them in writing after the first conversation, based on scope.

NEXT STEP

Bring us the system to audit →

Tell us what it does, what it costs and what worries you. We reply with what we'd need access to.

If the audit concludes it's better to start over, the next step is a → Operational product build