How to Build a Claude Agent: Tools, State, and Workflows
Learn how to build a Claude agent that turns repeatable reasoning, tool calls, and multi-step workflows into a dependable app with state, guardrails, and governed access.

What a Claude agent is
A Claude agent is a program that gives Claude one defined job, a set of typed tools, and a loop that keeps calling the Messages API until the job is done. Your code runs every tool call, stores state outside the conversation, and enforces permissions. Claude decides which tool to call and when the work is finished. The model supplies judgment. Your application supplies everything else.
Key takeaways - A Claude agent combines the Messages API, typed tools, an execution loop, application-managed state, and policy checks. - Claude handles judgment. Your application executes tools, stores state, and controls side effects. - Move repeatable steps into a governed Major app so the agent reasons once and runs the software on later jobs.
A chatbot answers and forgets. An agent takes actions across several turns, and actions can go wrong. That is why most of this article covers the parts around the model.
Some jobs should never become agents. If the steps are fixed and the inputs are structured, write ordinary code. Reach for Claude when a step means reading unstructured input and making a call a rule cannot encode, like deciding whether a diff touches authentication logic.
Choose one job and define completion
Narrow agents finish. General ones wander. "Triage every new pull request and post a review summary to Slack" is a job. "Help the engineering team" is a wish.
Write down what done means before you write a prompt. For PR triage, done means a findings record exists, a Slack message went out, and nothing was written to GitHub without approval. That definition becomes your system prompt and your test cases.
Define a small typed tool surface
Every tool is a function your code runs when Claude asks for it. Anthropic's tool use documentation defines three parts: a name, a description Claude reads to decide when to call it, and an input_schema in JSON Schema with its required fields. Claude picks tools by description, so five precise tools beat twenty overlapping ones.
{ "name": "get_pr_diff", "description": "Fetch the unified diff and CI test output for one pull request. Read-only. Call this before writing any review finding.", "input_schema": { "type": "object", "properties": { "repo": {"type": "string", "description": "owner/name, e.g. acme/billing-api"}, "pr_number": {"type": "integer"} }, "required": ["repo", "pr_number"] }, "strict": true}
Setting strict: true makes Claude's calls match your schema exactly. Split read tools from write tools, and name the write tools so the risk is visible. request_merge_approval tells a reviewer far more than update_pr.
Implement the Messages API tool loop
Every Claude agent reduces to the same four steps:
- Send a request to the Messages API with the system prompt, the
toolsarray, and the conversation so far. - Read
stop_reason. If it istool_use, the response holds one or moretool_useblocks, each with anid, aname, and aninput. - Run each call in your code and build one
tool_resultblock per call, matched bytool_use_id. Mark failures withis_error: trueand a short explanation. - Append Claude's response as an assistant message and your results as the next user message, then send again. Repeat until
stop_reasonisend_turn.
Two details break naive loops.
Claude can request several tools in one response. Run independent reads concurrently and side-effecting calls in order, then return every tool_result together in the next user message, before any text. If you skip a call because an earlier one failed, still return a result for it with is_error set. Anthropic's tool use overview covers this formatting, and the API rejects requests where a tool_use id has no matching result.
Non-success stop reasons need their own branches. max_tokens means output was cut off, and an incomplete tool_use block should be retried with a higher limit. pause_turn means a server-tool loop hit its iteration cap, and you resume by sending the assistant content back. refusal arrives as an HTTP 200 with stop_details naming the policy category. A loop that only checks for tool_use will log a refusal as a finished job. Cap total iterations too.
Persist state outside the context window
The messages array is Claude's memory for one run. When the process ends, it is gone. Replaying yesterday's transcript into today's prompt costs more every day and still loses detail.
| State | Where it lives | What it is for | |---|---|---| | Transcript | The messages array, in memory | Reasoning inside one run. Summarize or discard afterward. | | Database | Postgres or a managed app database | Findings, statuses, which PRs are done. The next run queries it. | | Files | Object storage | Diffs, test logs, and reports too large for a row. | | Audit log | Append-only table | Every tool call, its input and result, and who approved it. |
Give Claude a read tool over the database, such as get_prior_findings, instead of replaying history. The next run starts with a query.
Connect Anthropic to Slack, Notion, Gmail, or custom apps
The Anthropic connector on Pipedream authenticates with an API key and exposes a Chat action, with no triggers of its own. Events have to come from somewhere else: a GitHub pull request webhook, a new Gmail message, or a Slack slash command. Output goes to Slack for notifications and approvals, Notion for decision records, Gmail for external replies, and custom apps for anything your SaaS tools do not model. The Google Sheets automation pattern has the same shape with a different destination. If different steps deserve different Claude models, LLM routing covers that choice.
Each connector becomes a tool with its own credential. Scope each one to the minimum the job needs.
Add approvals, permissions, retries, and logs
The approval gate is where an agent earns trust. A tool that writes to a system of record should never execute directly. It creates a pending request, posts it to Slack with approve and reject buttons, and returns a tool_result telling Claude the action is waiting on a person. The write runs only after someone clicks approve, and the click is logged with their identity.
Permissions belong in credentials. Never put an API key or secret in a system prompt. "Do not merge" in a prompt is a suggestion to a model. A GitHub token without merge scope is a fact. Retries belong in tool code with backoff and idempotency keys, so a retried Slack post does not land twice. Every call goes to the audit table as it happens.
Claude remains part of the job. Deciding whether a diff is risky is still Claude's call. The guardrails decide what happens next.
What this approach does not cover
This guide covers a tool-using application agent with application-managed state and approvals. It does not specify a production deployment model, model-evaluation suite, latency budget, or a universal choice of Claude model. Test those decisions against your data, privacy requirements, failure modes, and workload before launch.
Build this in Major
Here is the weekly code-review agent end to end. A new GitHub pull request fires the trigger. Claude reads the diff and CI output through get_pr_diff and decides which findings matter: a missing null check in a payment handler, a test that asserts nothing. The app stores those findings in its own database and posts a summary to Slack with an approval request. When a reviewer approves, the decision and the reviewer's name are written to Notion. The GitHub, Slack, Notion, and Anthropic credentials are scoped to read plus comment. Merge stays with a human.
The review checklist is where the cost structure changes. On early runs, Claude works out that every PR gets the same checks. Migrations need a down step. New endpoints need auth middleware. Changed modules need tests. Re-deriving that checklist on every pull request is how agent cost climbs with usage. On Major, the agent builds those checks into a deterministic app that runs the same way on every PR, keeps its history in a managed database, and writes its own logs. Claude keeps the calls that need judgment, like whether a finding is serious enough to block. Reason once. Run forever.
Hand-building the loop, database, approval flow, and audit table is reasonable for one prototype. The maintenance burden appears when each agent has its own credential handling and partial log table. Major is the enterprise platform where agents build the software they run on, so SSO, scoped credentials, permissions, and audit apply to every app an agent builds from the first deploy. Token spend is front-loaded while the agent figures out the checklist, then flattens as the app carries the repeatable work. The same structure fits multi-step agent workflows in project management.
The PR triage agent is a good first build because the job is narrow, the checklist repeats, and the approval point is obvious. Build your Claude code-review agent on Major.
Related articles
- Building AI agents: a practical guide to production-ready systems
- AI workflow automation: from repeated prompts to governed apps
- LLM routing: how it works and what the benchmarks show
Related articles
Frequently asked questions
- What is a Claude agent?
- A Claude agent is Claude plus four things your application provides: typed tools it can call, a loop that keeps sending requests until the job is finished, state stored outside the conversation, and a policy for what it may do. Claude handles the judgment. Your code runs each tool, keeps records, and enforces permissions and approvals.
- How does Claude use tools?
- You send a Messages API request that includes tool definitions. When Claude wants a tool, the response has stop_reason tool_use and one or more tool_use blocks. Your code runs each call and sends back matching tool_result blocks in the next user message. Claude continues from there until it returns end_turn with a final answer.
- Can Claude remember across sessions?
- Not by itself. The messages array only lasts for one run. To remember across sessions, your application stores findings, statuses, and files in a database or object storage, then gives Claude a read tool to query them at the start of the next run. That is cheaper and more reliable than replaying old transcripts.
- How do you secure a Claude agent?
- Give each tool its own credential scoped to the minimum the job needs, such as read plus comment on GitHub with no merge rights. Keep secrets out of prompts. Route every write to a system of record through a human approval step, and log each tool call with its input, result, and approver in an append-only audit table.