3
4 Comments

Your AI Audit Log Is Probably Missing the One Thing Auditors Ask First

We keep seeing the same pattern in teams adopting multiple AI models.

They have provider dashboards. They have application logs. They may even have a security policy that says “do not send confidential data to public models.”

Then someone asks a simple question: who made this request, with which credential, under which policy, and what happened to the data?

The answer is often spread across four systems, two spreadsheets, and one person’s memory.


The audit target is a call chain

An AI request is rarely just a request from an app to a model. It can pass through a user or service account, an SDK, a proxy, a provider, a retrieval system, and an agent tool.

That creates a different security problem from traditional API monitoring. A successful status code does not prove that the request was authorized or that the right data policy was applied.

We think every request should be traceable across five questions:

  1. Who or what initiated it?
  2. Which project and environment was involved?
  3. Which virtual credential and policy version were active?
  4. Which logical model was requested, and which actual provider served it?
  5. Did a detector fire, and what action followed?

Start with a shared event vocabulary

Provider dashboards are useful, but they do not use the same identity, model, usage, or error vocabulary. A multi-model team needs a normalized event record.

Our minimum set includes identity, project, environment, application, logical model, actual model, provider, credential alias, policy version, data action, status, usage, cost, and trace ID.

The logical/actual model split is important. An application can ask for a capability such as “code review” while routing rules choose a concrete model based on availability, budget, or data restrictions. An audit needs to show both the intent and the final route.


Credentials are part of the audit story

Shared, long-lived API keys create an attribution problem before they create a billing problem. If five services use one key, revoking access for one service may affect all five. If a key leaks, the audit trail often ends at the provider account.

Scoped virtual credentials give the team a more useful unit of control. Record issuance, rotation, renewal, revocation, allowed models, rate limits, quotas, and the time at which a policy became effective.

The effective time matters. A current configuration cannot explain an event that happened three weeks ago after several policy changes.


You do not need to store every prompt in plain text

Many teams hesitate to build audit logging because they do not want to create another sensitive data store. That is a valid concern.

For many requests, the audit record can store the data classification, detection rule, redaction or blocking action, policy version, request owner, and outcome. Keep raw content only where the business and retention policy require it, and record who can view it.

The useful question is not “did we save every prompt?” It is “can we prove that the right control ran for this request?”


Test the chain with one random request

The most practical readiness test is simple: pick one real request at random and trace it from caller to credential, policy, route, data action, usage event, and incident outcome.

If a step requires a manual lookup or someone’s memory, that is the next engineering task. This test also works before a production rollout: validate the event model with a low-risk project, then expand it.


What we are building around this problem

At AiKey, we focus on the layer between enterprise identity and model execution. A local encrypted vault and proxy can keep credentials out of source code. Scoped virtual credentials, policy binding, routing, budgets, and audit events can connect one request to a user, project, environment, and logical model.

The goal is modest and practical: make the evidence form when the request happens, so an audit does not become a last-minute reconstruction exercise.

If this is a problem your team is working through, you can learn more about AiKey at https://aikeylabs.com/zh/i/ih40. For enterprise inquiries, contact aikeyfounder@gmail.com.

on September 28, 2026
  1. 1

    Yep, dashboards and app logs often show when you called a model and what it cost. The harder question is: what did the model see, what did it return, and what happened next?

    A useful audit trail should capture:

    • Who made the request and when.
    • The model version and settings.
    • The prompts, conversation history, and retrieved context.
    • The original response and any changes made afterward.
    • Tool calls, results, and errors.
    • Human approvals or rejections, including the reason.
    • Token usage, cost, latency, and retries.

    A simple check is to pick a bad response and see whether your logs explain how it happened. You may not reproduce the exact answer, but you should be able to reconstruct the steps.

    Keep those details linked under one request ID, version your prompts, and protect sensitive information with appropriate access and retention controls. Hashes can verify that stored content hasn’t changed, but you still need access to the content itself to investigate.

    That’s what makes logs useful when something goes wrong: enough evidence to understand the issue and fix it.

  2. 1

    Nice, this makes a lot of sense. What's been the most surprising part of it so far?

  3. 1

    The audit trail needs the decision context, not just which model answered. I would log the initiator, policy version, tools or data touched, outcome, and any human approval gate, while redacting raw prompts when possible. What is the one field you wish you had logged on day one?

  4. 1

    The policy version point is exactly what we found missing when we checked our own setup. Our AI agent calls the same write handlers as the UI and acts as the signed-in user. Every event it writes carries that user's id and a request id built from the agent run and the tool call, so from any change you can get back to the conversation and the tool call behind it.

    We can't answer why the write was allowed without asking. Users can set an action to "always allow". A write the user confirmed points back to the proposal they approved. An auto-approved write has nothing like that: we don't record which rule let it through, and the user can change that rule the next day. We also don't record which model made the call.

    Do you store a policy version on each request and keep the old versions, or a copy of the rule as it was applied?