Skip to content

AI Agent Reliability Audit

I check your AI agents before your customers do.

Putting AI agents near money, customers or production systems? There are a handful of predictable ways they fail — a checkout that confirms on error, an injection that exfiltrates data, a webhook that runs twice. I force each of those to happen once, under control, so it doesn’t happen to your customers. The same discipline I run on my own agent fleet in production.

The accountable human holding the trust boundary — not a scanner, not an agency.

Who it’s for

For teams that gave an AI agent real responsibility

Agents near money

Checkout, refunds, payouts, invoices — anywhere an agent can confirm or move value.

Agents near customers

Support, email, messages that go out in your name — where an injection or a mistake becomes the customer’s problem.

Agents near production

Agents with tools that write to databases, call APIs or run commands in your systems.

The method

The 7 lenses

An agent system fails in a small number of predictable ways. Each lens forces one of them to happen — under control, with evidence.

  1. 01

    Surface inventory

    Every way in: entrypoints, webhooks, crons, and which tools each agent is allowed to call.

    How it’s tested: Derived from code, never an old diagram. Every unauthenticated entrypoint and over-broad tool grant is flagged.

  2. 02

    Injection paths

    Every external string that reaches a prompt — webhooks, email, RAG chunks, fetched web pages.

    How it’s tested: Adversarial injections that try to exfiltrate data, trigger an unapproved action, or escalate the agent’s permissions.

  3. 03

    Money paths

    Does the server trust an amount or success flag from the client? Are failures honest? Can everything be reconciled?

    How it’s tested: I force the unhappy paths: declined card, dropped connection, spoofed "paid". Money paths must fail closed.

  4. 04

    Idempotency & retries

    Every webhook and job runs at least once. Are there idempotency keys — and are they backed by a DB unique index?

    How it’s tested: I redeliver the same webhook twice and run concurrent jobs. Any double charges, prints or rows?

  5. 05

    Autonomy boundaries

    What may the agent finish alone, and what must stop for a human? Where are the approval gates?

    How it’s tested: I hunt the outward action that ships without a gate — and gates placed after the irreversible step.

  6. 06

    Limits

    Cost caps, daily call caps, rate buckets, loop guards, timeouts — what stops a runaway agent.

    How it’s tested: I try to blow past the cap and check the runner refuses to start — not just logs a warning and continues.

  7. 07

    Observability

    When an agent misbehaves at 02:00 — can you tell what it did, why, and whether it happened before?

    How it’s tested: I trigger a controlled failure and ask the system to explain it from its logs alone. If it can’t, that’s a finding.

What you receive

Evidence, not opinions

A finding is never "you should improve X". It’s a reproducible failure scenario, the smallest fix that closes the root cause, and the step that proves it holds.

A written report

An exec summary, a scorecard per lens, and findings ranked by blast radius — what breaks and how far it spreads.

Concrete failure scenarios

Each finding: the exact input, the outcome it produced, the root cause, the fix and the verification step.

A 30-day fix plan

Ordered by risk-per-hour — the biggest risk per unit of effort first.

A re-test

Within 30 days I re-run every probe that failed. A finding closes when its verification step passes — not when someone says it’s fixed.

Pricing

Two ways to start

Agent Health Check

A fast, honest temperature check: how dangerous is your agent setup right now?

9 995 krexcl. VAT

About 1 day · fixed price

  • A surface map of your agents’ entrypoints, tools and webhooks
  • The top 5 risks, ranked by blast radius
  • A prioritised shortlist of what to fix first
  • Fee credited toward a full audit if you upgrade within 30 days

9,995 kr fixed price · excl. VAT

Full audit

AI Agent Reliability Audit

The full 7-lens audit of one agent system — report, fix plan and re-test.

34 995 krexcl. VAT

About 1 week · fixed price

  • Full review across all 7 lenses, with adversarial probing
  • A written report: scorecard, findings ranked by blast radius
  • Each finding: a concrete failure scenario (input→outcome), fix and verification step
  • A 30-day fix plan, ordered by risk-per-hour
  • Re-test within 30 days: I re-run every probe that failed

34,995 kr fixed price · excl. VAT

Why me

I run this discipline on my own system, every day

I’m not a scanner and not an agency. I’m one person in Sweden who builds and runs an agent fleet in production — agents near money, print-on-demand and human-in-the-loop approvals. This quarter I hardened it: prompt-injection defense, HMAC on webhooks, idempotency end to end, per-client cost caps, money paths that fail closed. The audit is that practice turned outward. I ran an honest self-audit of my own platform with the exact same template — it’s there to read.

Honest scope

What an audit is not

I sell evidence, not false comfort. Here are the boundaries, stated up front.

  • Not a certificate. This isn’t ISO 27001, SOC 2 or a formal compliance stamp — it’s a technical review of how your agents fail.
  • No guarantee the system is "secure". I find and prove the failures the method covers; I don’t promise nothing else exists.
  • I don’t pentest your third-party vendors. I audit how you use them, not their internals.
  • I don’t fire live probes at production money paths without written scope. Work runs against staging/test mode.
  • I don’t fix everything for you in the price. The audit delivers findings, proposed fixes and a re-test; the implementation itself we can agree on separately.

Frequently asked

How long does an audit take?

A full Reliability Audit takes about one focused week from kickoff to report, plus the re-test within 30 days. The Health Check is about a day.

What do you need from us?

Read access to the agent system’s code, a staging or test environment to run probes against, and a contact who can answer how the flows are meant to work. No production access is required for the core work.

Do you touch production or real money?

No, not without written scope. Adversarial probes run against staging or test mode. Money paths are tested with provider test keys, never live charges unless we’ve agreed on it.

What frameworks does the method build on?

The 7 lenses are my own practice, hardened on my own agent fleet, but they map well onto OWASP’s Top 10 for LLM and agentic apps. The difference is I force each failure to actually happen — not just check a list.

Do you fix what you find?

The audit delivers findings, a proposed fix per finding and a re-test. If you want me to implement the fixes we agree on that separately — but many teams fix it themselves with the plan in hand.

Is the report confidential?

Yes. The report names real failure modes and is a map of your attack surface if it leaks — it’s handled accordingly and shared only with you.

How is this different from the Health Check?

The Health Check is a rapid triage of about a day: a surface map plus the top risks, without deep adversarial probing or a formal re-test. It’s the door-opener — and its fee is credited toward a full audit if you upgrade within 30 days.

Got agents near money or customers?

Book an audit and we’ll go through your agent system, what’s in scope and which package fits. No sales pitch.

Book an audit