LLM Observability: What to Instrument When the Work Is a Black Box

Tracing a model call tells you what the model did, not why the work was right. The vendors selling LLM observability skip the standard that makes traces portable, and skip the harder question: how much of your system needs watching in the first place.

Rahul Ramakrishnan
llm-observability-hero.png

What a CIO or CISO actually needs to know

A trace tells you what the model did. It does not tell you whether anything in your organization authorized it to do that.

The budget request usually arrives framed as tooling. Someone owns an LLM feature that shipped last quarter, security has started asking questions, and the proposed answer is a platform. The question underneath is narrower and harder. When this system produces a wrong answer, sends the wrong email, or writes to the wrong record, can you reconstruct what happened, in what order, using whose credentials, and can you show that reconstruction to someone who was not in the room?

Most teams instrument for debugging and then discover they were asked for accountability. Those are different requirements. Debugging wants enough signal to reproduce a fault while the incident is warm. Accountability wants a record that survives a year, attributes an action to an identity, and holds up when the person reading it is skeptical. A tracing backend with fifteen days of retention answers the first and not the second.

That distinction is where most LLM observability programs quietly fail. The instrumentation is real, the dashboards are populated, and the audit question still has no answer, because nothing in the trace records authority. Observability tells you what happened. Governance decides what was permitted, and it needs control at the point of action rather than a report afterward. If you arrived here from a policy mandate, the gap between the two is the whole problem: see from policy to implementation.

What LLM observability covers that APM does not

Application performance monitoring assumes failure announces itself. A request throws, a status code changes, a queue backs up. The signal that something went wrong is generated by the thing that went wrong.

Model-backed work breaks that assumption. A wrong answer returns HTTP 200. A hallucinated invoice number is well-formed JSON. Nothing in the transport layer knows the difference between a correct summary and a confident fabrication, which means correctness has to be measured by a separate system rather than caught by an error handler.

IBM's page on the topic groups the measures into system operation, resource consumption, and model behavior, and dates itself honestly (published 25 February 2025, updated 22 June 2026, ibm.com). That grouping is right as far as it goes. It undercounts one thing: provenance. Knowing that a call was slow and expensive and probably wrong still leaves you unable to say which prompt version, which tool inventory, and which retrieved documents produced it.

| Signal | In APM | In model-backed work | | --- | --- | --- | | Failure detection | Exceptions and status codes | Output can be wrong while the call succeeds; needs a separate evaluation pass | | Latency | Request duration | Duration plus time to first token and time per output chunk, since streaming changes what users perceive | | Cost | Indirect, via infrastructure | Direct and per-request, metered in input, output, reasoning, and cached tokens | | Correctness | Not a signal | The primary signal, and it has no error code | | Provenance | Code version from the deploy | Prompt version, model version, tool inventory, and retrieved context, none of which are implied by the deploy | | Input identity | Structured parameters | Free text, which means the input is also a payload with a redaction problem |

Langfuse's definitional page covers the conceptual half of this cleanly and lists runtime tracing, evaluator-based quality measurement, user-intent analysis, and production feedback loops (langfuse.com). It never mentions OpenTelemetry, which is a strange omission on a page defining the category, and a useful signal about where the market's attention is.

The standard almost nobody is writing about

Search the phrase and read the first ten results. Datadog's page mentions OpenTelemetry as an instrumentation option and does not name the GenAI semantic conventions. Langfuse does not mention OTel at all. The vendor listicle ranking third does not either. IBM gestures at it once, in a closing paragraph about connecting to broader monitoring platforms.

Meanwhile there is a specification, it is being actively worked on, and it decides whether the traces you emit this year are readable by whatever you buy next year.

What the GenAI semantic conventions actually specify

The conventions define the span shape for a model call. The naming rule is mechanical: "Span name SHOULD be {gen_ai.operation.name} {gen_ai.request.model}" (gen-ai-spans.md). Two attributes are Required, gen_ai.operation.name and gen_ai.provider.name. Everything else is conditionally required, recommended, or opt-in.

gen_ai.operation.name carries 18 allowed values, and the list tells you how far the spec's ambition has grown past chat completion: chat, create_agent, create_memory, create_memory_store, delete_memory, delete_memory_store, embeddings, execute_tool, fetch_response, generate_content, invoke_agent, invoke_workflow, plan, retrieval, search_memory, text_completion, update_memory, upsert_memory. There is a value for planning and a value for invoking a workflow. The working group is modeling agents, not just calls.

| Attribute | Requirement level | What it tells you | | --- | --- | --- | | gen_ai.operation.name | Required | Which of the 18 operations this span represents | | gen_ai.provider.name | Required | Which provider served the call | | gen_ai.request.model | Conditionally required | The model you asked for | | gen_ai.response.model | Recommended | The model that actually answered, which is not always the same | | gen_ai.conversation.id | Conditionally required | The correlation key that stitches a multi-turn interaction together | | error.type | Conditionally required | Failure classification when the call fails outright | | gen_ai.response.finish_reasons | Recommended | Why generation stopped, including truncation you would otherwise miss | | gen_ai.usage.input_tokens, gen_ai.usage.output_tokens | Recommended | Billable volume in each direction | | gen_ai.usage.reasoning.output_tokens | Recommended | Tokens spent on reasoning the user never sees | | gen_ai.usage.cache_read.input_tokens, gen_ai.usage.cache_write.input_tokens | Recommended | Whether prompt caching is working, separated from raw volume | | gen_ai.input.messages, gen_ai.output.messages | Opt-in | The content itself, off by default for a reason |

The metrics document is separate and defines a matching set, including gen_ai.client.token.usage, gen_ai.client.operation.duration, gen_ai.client.operation.time_to_first_chunk, gen_ai.invoke_agent.inference_calls, and gen_ai.invoke_agent.tool_calls (gen-ai-metrics.md). Counting inference calls and tool calls per agent invocation is the metric most home-grown instrumentation forgets, and it is the one that tells you an agent is looping.

Where the spec now lives, and what Development status means

Two facts that most writing on this topic has not caught up with.

First, the conventions have moved. The gen_ai path on opentelemetry.io is no longer maintained and now points readers to the GenAI semantic conventions repository. The most-linked URL in the ecosystem is a stub. If your internal design doc cites it, the citation is stale.

Second, and this is the part to say plainly: the GenAI semantic conventions are at Development status, not Stable. OpenTelemetry's own versioning policy is explicit about what that means. "While signals are in development, breaking changes and performance issues MAY occur," and "Long-term dependencies SHOULD NOT be taken against signals in Development" (opentelemetry.io). Both the spans and metrics documents carry the Development badge today.

There is also no rendered, versioned documentation site. The human-readable version is the markdown in the repo's docs directory, and the README still lists the schema URL as TODO. Anyone telling you this is a settled standard has not opened it.

What that changes in practice is less than it sounds. When other convention families migrated toward stability, OpenTelemetry shipped a dual-emission mechanism, OTEL_SEMCONV_STABILITY_OPT_IN, which takes category values like http or http/dup so instrumentation can emit old and new attributes together during a transition (http-migration). The GenAI docs do not currently document a gen_ai value for it. Plan for a rename cycle rather than a rewrite, and keep your attribute mapping in one place so a rename is a config change instead of a project.

Where prompts and responses belong

In 2024 OpenTelemetry's own blog recorded the LLM working group's guidance to capture prompt and response detail "on events instead of span attributes," on the grounds that backends struggle with very large payloads (opentelemetry.io/blog).

The current spec is looser than that summary. It defines gen_ai.input.messages and gen_ai.output.messages as opt-in attributes and permits content on spans or events, preferring structured form and allowing JSON strings where structured attributes are unsupported. The payload concern that produced the original guidance has not gone away. A conversation history attached to every span multiplies your storage bill by the length of the conversation, and it puts user content into a system whose access controls were designed for stack traces. Treat content capture as a decision with a retention policy and a redaction step attached, not as a checkbox. That is where the agent threat model stops being abstract.

What good looks like in practice

A baseline that clears an audit conversation, in order:

  1. Emit GenAI semantic-convention attributes from day one, starting with the two Required fields and the token-usage family. Naming is cheap now and expensive to retrofit across a year of stored traces.
  2. Set gen_ai.conversation.id on every span in a multi-turn interaction, and propagate your own business correlation id alongside it. Without both, a trace is a call and not an episode.
  3. Record the version of everything that shapes output: prompt version, model version, tool inventory, retrieval index version. gen_ai.response.model catches provider-side substitution that gen_ai.request.model hides.
  4. Turn token usage into an operational metric with an owner and a threshold, not a line item someone reads at month end. The attributes separate input, output, reasoning, and cached tokens for exactly this reason, and the follow-through belongs with token cost as an operational metric.
  5. Decide content capture deliberately. Redact at the instrumentation boundary rather than in the backend, sample full content at a rate you can defend, and keep structured metadata at 100 percent so the shape of an interaction survives even when its text does not.
  6. Link the model span to the side effect it caused. The row written, the message sent, the ticket closed, with the credential that authorized it. This is the step almost nobody does, and it is the only one that turns a trace into an audit record.
  7. Evaluate output on a schedule, not only during incidents. Correctness has no error code, so the absence of alerts means nothing at all.

Step six is worth dwelling on. Everything else on this list is available from an instrumentation library. Attribution of a side effect to an identity is an architectural property of the system taking the action, and no tracer can add it after the fact.

Where the market is getting this wrong

Be specific about this, because the pattern is consistent.

Datadog's page ranks near the top for a definitional query and opens with "Ship AI agents faster, with confidence" under a title tag that still reads "Agent Observability | LLM Observability" (checked 8 September 2026, datadoghq.com). A reader looking for a definition gets a hero CTA, and the one conceptual answer on the page sits in an FAQ near the bottom. The figures it reports, including a percentage improvement in MTTR and a percentage reduction in token usage per task, are self-reported with no methodology, no sample, and no date. Those numbers may well be real. They are not checkable, which is the same thing as unusable in a procurement review.

The vendor listicle ranking third is worse in a more interesting way. It ranks itself first, states evaluation criteria that map neatly onto its own feature set, carries no publication date on a page whose entire premise is currency, and reports a competitor's pricing in a completely different shape from what that competitor advertises on its own page. I am deliberately not printing either set of figures side by side, because one is uncited and both go stale. The methodological point stands on its own: an undated comparison with uncited pricing is not research.

And then the silence on the standard, which is the real gap. Four of the highest-ranking pages defining this category either ignore the specification or mention it once in passing.

Here is the position that follows, and it is uncomfortable given who ranks for this term. Most teams do not need an LLM-specific observability vendor. They need GenAI semantic conventions emitted into the OTel stack they already run, with the six steps above actually done.

That position has real limits, and pretending otherwise would be dishonest. Raw OpenTelemetry gives you traces and metrics. It does not give you evaluation workflows, an evaluator library, prompt-level analytics, dataset management, or a place for a non-engineer to review flagged outputs. Those are genuine capabilities and a genuine reason to buy a platform. The argument is about sequencing rather than abstinence. Emit the conventions first, because that keeps the decision reversible, and buy the evaluation layer when you have a measurable quality problem rather than in anticipation of one.

The question the tooling debate skips

Every page ranking for this term treats the amount of observation you need as a fixed quantity and competes on tools to observe it. That skips the prior question. How much of your system has to be watched from the outside in the first place?

If a workflow re-reasons its way through the same six steps on every run, then every run is a fresh black box and tracing is your only recourse. You are reconstructing intent from spans after the fact, which is forensics. It works, at a cost that scales with how often the workflow runs and how much of it is model-driven.

If the repeatable part of that workflow has been moved into deployed code, that part is not a black box. It has its own logs, its own database rows, its own audit trail, and a diff history that says when its behavior changed and who changed it. You do not point a tracer at it to find out what it does. You read it. Tracing is then reserved for the steps that genuinely required judgment.

This does not remove the need for LLM observability, and any vendor telling you it does is selling something. The model still runs, still needs tracing, still needs evaluation. The claim is narrower. Observability effort scales with how much of your system re-reasons, and that quantity is a design decision rather than a fact of nature. It is decided when you choose how agentic workflows are structured, long before anyone picks a backend.

This argument is about the model layer. The related problem of instrumenting a multi-step agent, where the hard part is capturing decision context and tool inventory at the moment of action, is observability for agents specifically and deserves its own treatment. Evaluation methodology is a third problem, harder than either, and nothing here addresses it.

What we're doing about this at Major

We think observability is framed as a tooling problem when it is substantially an architecture problem. Every step that re-reasons through a model on each run is a step you can only understand afterward, by reconstruction. Most teams accept that surface as fixed and then buy tools proportional to its size.

Major is the enterprise platform where agents build the software they run on, and this is the reason the design matters here. When an agent on Major works out how to handle a repeatable part of a task, it builds an app for that part and runs the app instead of reasoning through the step again. The app holds its own data in a managed database and writes its own logs, so that portion of the work is inspectable because it is code, not because a tracer was pointed at it. Scoped credentials and audit logging apply where the agent acts, which means an action can be attributed to an identity rather than merely observed in a span.

The honest scope: the model still runs, and the model layer still needs tracing and evaluation with the semantic conventions above. What changes is the size of the surface requiring black-box observation. It shrinks toward the judgment calls instead of covering the whole workflow, and the cost of that shift is front-loaded rather than recurring. Reason once. Run forever.

If step six on that baseline list is the one your current stack cannot answer, the attribution problem is worth looking at from the architecture side: see how Major ties an agent's actions to scoped credentials and its own audit log.

Related articles

Frequently asked questions

What is observability in LLMs?
LLM observability is the practice of instrumenting a model-backed application so its behavior can be inspected after the fact. It combines three things: traces of each model call and tool invocation, metrics such as latency and token usage, and evaluation of output quality, which has no error code and must be measured separately.
How is LLM observability different from traditional APM?
APM assumes failure announces itself through an exception or a status code. A model can return a wrong, fabricated, or unsafe answer with HTTP 200 and no anomaly anywhere in the transport layer. So correctness becomes a signal you have to measure rather than catch, and cost per request and prompt provenance become first-class telemetry instead of infrastructure details.
What should I instrument for an LLM application?
Follow OpenTelemetry's GenAI semantic conventions. Emit the required gen_ai.operation.name and gen_ai.provider.name, plus the requested and responding model, gen_ai.conversation.id, finish reasons, and the token-usage family covering input, output, reasoning, and cached tokens. Add latency and time to first chunk. Then link each model span to the business action it caused and the credential that authorized it.
Does OpenTelemetry support LLM observability?
Yes, through the GenAI semantic conventions, with two caveats worth knowing. The conventions are at Development status rather than Stable, so breaking changes are permitted and long-term dependencies are discouraged. They have also moved off opentelemetry.io to the semantic-conventions-genai repository, where the docs directory holds the human-readable version. There is no rendered, versioned documentation site yet.
What is the best LLM observability platform?
Judge candidates on four things rather than a ranking. Whether traces carry GenAI semantic-convention attributes instead of a proprietary schema, whether evaluation is a real workflow or a dashboard, what retention you get and what it costs at your volume, and whether a model span can be linked to the action it caused. Emitting the standard keeps that choice reversible.