Agent Observability: What Deterministic Apps Change
Tracing model calls matters, but it is only part of agent observability. See how deterministic apps make repeated execution inspectable and change what teams need to monitor.

What the CIO and CISO actually need to know
Agent observability, as used here, is the ability to reconstruct what an agent workflow did: which model calls were made, which tools were invoked, what state changed, under whose credentials, and how long each step took. Google Cloud defines it as "the methods for gaining insights into the internal state and behavior of software agents" and lists six things to watch: LLM interactions, tool usage, agent behavior and reasoning, performance, security and safety, and quality and evaluation.
That list is correct and incomplete. Every item on it describes telemetry. None of it says who was allowed to do what at the moment of action. A team can have flawless traces and still be unable to answer the question an auditor asks first: was this action permitted?
So we separate four things that vendor pages tend to blur.
| Concept | Question it answers | Output | |---|---|---| | Observability | What happened at runtime? | Traces, logs, metrics | | Evaluation | Was the output good? | Scores against criteria | | Deterministic testing | Does the code do what it should? | Pass or fail assertions | | Governance | Was the action allowed, and by whom? | Scoped credentials, approvals, audit events |
This article does not cover agent evaluation methodology or prompt-injection defense. Both deserve their own treatment. It also does not claim that deterministic apps remove the need for tracing. They change where you point it.
The argument
Most agent work is a mix of two kinds of steps. Some steps need judgment: classifying an ambiguous request, deciding whether a contract clause applies. Others are repeatable: routing a record, updating a status field, writing to a system of record. A single observability approach applied to both wastes effort on one and under-serves the other.
Our position is that observability should match the execution surface. Trace what the model decides. Log what the app executes. Correlate the two. Everything else follows from that.
Trace what the model decides
Model-driven steps are probabilistic, so they need the full apparatus. Langfuse describes a trace that records LLM calls with prompts, completions, token usage and cost, and tool calls that distinguish the tools available from the tools actually invoked. That distinction matters for audit. If you cannot see what the agent could have done, you cannot judge whether what it did was reasonable.
OpenTelemetry's GenAI work is defining semantic conventions for models, vector databases and agents, and describes AI agents as non-deterministic, which is why telemetry doubles as a feedback loop for quality. The agent conventions are still evolving. Verify the current state of the spec before you standardize on attribute names.
Datadog's agent observability product covers the same territory, tracing prompts, retrieval steps, tool calls and agent decisions, and correlating LLM spans with APM services and infrastructure signals. Teams that already run Datadog get real value from that correlation. We are not arguing against any of it.
Model traces carry fields such as model, prompt version, tool inventory, selected tool, arguments, latency, token counts, and the agent's stated reasoning where the model exposes it.
Log what the app executes
A deterministic app step is code. It takes a defined input, applies a rule, and produces a result. You do not need a reasoning trace to explain why a ticket in the billing category routed to the billing queue. You need a record that it did, with the inputs, the actor, the permission that allowed it, and the state before and after.
App logs carry fields such as execution ID, operation name, input reference, output or status, state transition, credential scope, timestamp, and error path. Because the app holds its own database, the state change is a row you can query, not a sentence in a context window.
This is the claim that belongs to Major specifically. An agent that learns a procedure and performs it through the model again on every run leaves you with a reasoning trace to interpret each time. An agent that builds an app for the repeatable part leaves you with code that ran the same way and a log of what it did. Reason once, run forever applies to observability too: the second run is easier to read than the first.
Observe the handoff
The riskiest point in a hybrid workflow is the seam between a model decision and an app invocation. A classification was made, an app was called, a record changed. If those three events live in three systems with three identifiers, the seam is invisible.
Use three IDs and propagate them everywhere:
agent run ID -> app execution ID -> business record ID(model trace) (app log + state) (ticket, invoice, account)
The agent run ID joins the model trace. The app execution ID joins the app log. The business record ID joins both to the thing the business cares about. Any one of them should lead you to the other two in a single query.
Evaluate judgment, test repeatable code
Evaluation and testing are different instruments. Langfuse's guide makes a useful split: human annotation for ambiguous cases, code evaluators for deterministic properties, and LLM-as-a-judge for semantic calls like groundedness. That maps cleanly onto our model. Evaluations score the model's judgments. Conventional tests and run assertions cover the app logic.
What changes with a deterministic app layer is the evaluation budget. If routing, status changes and field updates run in code with their own tests, your evaluation effort concentrates on the classification and the exceptions. You are not re-scoring a hundred fixed steps for every judgment call. Neither instrument substitutes for monitoring, and monitoring substitutes for neither.
Put privacy and retention in the design
Prompt and output telemetry contains whatever the agent saw: customer messages, account details, contract text. Telemetry that is not redacted is a second copy of your sensitive data with weaker access controls than the original.
Decide four things before you instrument. What gets redacted before it leaves the runtime. Who can read raw prompts versus summaries. How long each class of telemetry lives. Whether you can export it when a regulator or customer asks. Platform behavior varies on all four, so verify retention, redaction and export in the current documentation of whichever vendor you choose. Observability tools also do not prevent unsafe actions by themselves. They record them.
What good looks like in practice
Take a support-ticket triage workflow. This is an illustrative walkthrough, not a captured production trace.
A ticket arrives. The agent reads it and classifies intent. Some tickets are ambiguous. That is a judgment, so it gets a full model trace: prompt version, the classification, the tool inventory, latency, tokens. The agent then calls a routing app. The app applies deterministic rules, assigns the queue, sets the status, and writes a row to its own database. A refund above a threshold needs approval, so the app emits an approval-required event, a person approves, and the write lands under a scoped credential with an audit event.
The minimal observability map for that one workflow:
| Surface | Signals | Diagnostic question | Retention and privacy risk | Owner | |---|---|---|---|---| | Model reasoning | Prompts, completions, tool selection, tokens, latency | Why did it classify this ticket as billing? | High: raw customer text in prompts. Redact, restrict, short retention | AI engineering | | App execution | Execution ID, operation, state before and after, error path | Did routing run, and what changed? | Medium: record contents in state. Follow source system rules | Platform | | Handoff | Agent run ID, app execution ID, record ID | Which decision caused this change? | Low: identifiers only | SRE | | Action and approval | Credential scope, approver, timestamp, result | Was this write permitted, and by whom? | Medium: identity data. Long retention for audit | Security | | Quality | Eval scores, human review rate, drift | Is classification getting worse? | Medium: labeled samples contain real data | AI engineering |
An event schema checklist for every action the workflow takes:
- Timestamp.
- Actor and credential scope.
- Operation name.
- Input reference, not the raw input.
- Result or status.
- Error path or approval path taken.
Model traces get the prompt-and-reasoning fields. App logs get the six above plus the state transition. Both carry the shared IDs. Together they let you answer, for any record, which tools were available, what the agent decided, what side effects landed, under whose credentials, and whether the run can be repeated.
Where vendors are getting this wrong today
The observability platforms are good at what they do, and the teams building them have moved fast. Google, Datadog and Langfuse all cover traces, evaluations and monitoring, and Datadog and Google also list security and governance features. Our critique is narrower and concerns framing.
The framing treats the trace as the control plane. A trace records that an agent called a tool. It does not establish that the call was authorized, and in most designs it records the action after the fact. Google's own list puts "security and safety" as one of six monitoring dimensions, which is monitoring policy enforcement, not enforcing it. That is the right scope for an observability product. The mistake is buyers assuming it is also the authority boundary.
The second problem is uniform treatment. When every step in a workflow is traced as if it were a model decision, the deterministic steps generate telemetry volume without much diagnostic value. The architecture can go further: make the repeatable steps code, and they stop being spans you interpret and become executions you query.
What we're doing about it at Major
Our position is this. Trace the reasoning. Inspect the run. Govern the action. Three different jobs, and no single tool does all three.
Major is the enterprise platform where agents build the software they run on. When an agent works out a repeatable part of a workflow, it builds an app for that part. The app runs in deterministic code, keeps its state in a managed database, and writes its own logs. The model keeps the judgment calls. So the support-ticket routing above is an app with an execution record and queryable state, and the classification remains a model step that you trace and evaluate with the tooling you already trust.
Permissions apply where the agent acts, through scoped credentials and role-based access, with audit events on the action. That is the governance layer, and we do not ask a trace to do its job.
What this does not solve: model-driven paths stay probabilistic, and they still need traces, evaluations and redaction. The app layer does not replace your observability vendor. It shrinks the share of the workflow that depends on reading reasoning, and it gives the remainder a code-backed record. If you want the longer treatment, read why agent observability needs to cover execution, governing agents at the point of action, the AI agent security threat model, and LLM observability and black-box execution.
If you want to see what an agent-built app's execution record and audit trail look like for a workflow of your own, see how Major turns an agent's repeatable work into an inspectable app.
Related articles
Frequently asked questions
- What is agent observability?
- Agent observability is the practice of capturing runtime evidence of what an agent workflow did: model calls, tool operations, outcomes, timing, errors, and state changes. It diagnoses what happened. It differs from evaluation, which scores output quality, and from governance, which decides whether an action was permitted.
- What are examples of agent observability platforms?
- Datadog, Langfuse, and Google Cloud Observability each document agent tracing and monitoring, and OpenTelemetry is defining open semantic conventions for AI agent telemetry. Capabilities, retention, redaction, and export behavior change often, so verify them in each vendor's current documentation before you choose.
- What should an AI agent trace include?
- A trace should record model calls with prompt version, latency, and token usage; the tools available and the tool calls actually made; errors; and a shared identifier linking the agent run, any app execution, and the business record. Redact sensitive inputs and restrict access to raw prompts.
- How is agent observability different from AI evaluation?
- Observability captures runtime evidence so a team can diagnose what happened in a specific run. Evaluation scores outputs or behavior against defined criteria, using human review, code checks, or model judges. You need both: traces supply the cases, and evaluations tell you whether the behavior was acceptable.