NVIDIA releases open-source model Nemotron 3 Super›

Context engineering in practice: build a maintainable operating system for AI agents

context engineering · AI agents · prompt systems · evaluation · agent workflowReading time: 10 minPublished: 2026.09.16
Context engineering workstation showing AI agent prompt versions, evaluations, and rollback

Context engineering in practice: build a maintainable operating system for AI agents

Agent work often fails because its task, rules, examples and tool behavior sit in different places. The model fills gaps with a plausible guess. A longer prompt rarely fixes that; it carries the ambiguity into the next turn.

This playbook is for teams using Claude, Codex, or a similar agent that reads files and calls tools. Define the deliverable, expose unknowns, supply evidence at the right step, and use the same acceptance rule for every handoff.

Context engineering layered workflow from task brief through project rules, evidence, tool contracts, evaluation and maintenance

The five layers have different jobs. A task brief defines the immediate result. Project rules hold durable local conventions. Evidence establishes facts. A tool contract describes inputs, outputs and effects. Evaluation decides whether the handoff is usable.

A prompt is only one part of context

Fix the login flow omits a reproduction path, allowed change surface, security boundary and acceptance test. Anthropic’s public context engineering guide recommends giving an agent enough relevant material for its current step instead of loading everything.

Material Question Home Failure mode
Task brief What must ship now? Issue or task file Improve this
Project rules How does this project work? AGENTS.md or CLAUDE.md Ten-page manual
Skill How is one repeatable job done? skills/name/SKILL.md Loading every Skill
Reference What is true or correct? Tests, types, fixtures, specs Prose instead of evidence
Tool contract What may the tool do? Tool docs or wrapper Name-only description

Each rule needs one authoritative home. A release requirement belongs in project rules. Today’s release angle belongs in the task brief.

Audit context before editing

Ask for a small inventory before the agent reads the repository:

Do not change files yet.
1. Find relevant rules, code, tests, data and examples.
2. List path, purpose, last modification and stale-risk.
3. Flag conflicts, unverifiable claims and missing facts.
4. Propose the smallest plan: what to read, change and test.
When evidence is missing, ask. Do not invent it.

This catches conflicts a model should not decide: prose says userId while a fixture uses id; a root rule preserves reports while a workflow clears its output directory; an old screenshot is mistaken for a current interface.

Keep AGENTS.md light; put procedure in Skills

AGENTS.md is a map, not a manual. Keep startup commands, ownership, secret handling, destructive boundaries, test expectations and output locations.

# Project rules
- Web app is apps/web; API contracts are packages/contracts.
- Never put secrets in fixtures, logs or screenshots.
- API contract changes need a fixture and one integration test.
- Run pnpm test:affected before handoff.
- Reports go to artifacts/date; do not alter earlier reports.
- If rules conflict, stop and quote both paths.

Anything a repository can reveal on its own is usually noise. Long procedures belong in scoped Skills. A release-review Skill can inspect metadata, image paths and links only when a release is being prepared. Claude Code Skills are one example of loading a procedure on demand.

Task briefs need knowns, unknowns and acceptance

# Task: add export status
## Outcome
After export, show queued, running, completed or failed.

## Known facts
- POST /exports returns job_id.
- Status types: packages/contracts/export.ts.
- Polling utility: apps/web/lib/poll.ts.

## Unknowns to resolve
- Can failed jobs be retried?
- How long is status retained?
- Is analytics required? Inspect existing code first.

## Constraints
- Do not change authentication middleware.
- Add no dependency.
- Do not log export contents in the browser.

## Acceptance
- Cover four states and one timeout.
- Use existing fixtures.
- Report changed files, commands run and open questions.

Explicit unknowns prevent unspecified from becoming free to improvise. If evidence cannot answer a question, the agent should stop and ask.

Connect evidence, tools and evaluation

A current type definition, JSON fixture or passing integration test is more useful than a paragraph saying how an API probably behaves. Scope references:

packages/contracts/export.ts establishes API field names only.
docs/ui/export-status.png establishes spacing only, not API behavior.
If tests conflict with prose, follow tests and types, then report the conflict.

Tool descriptions are interfaces. State input, output, side effect, permission and failure behavior:

tool: get_order
input: order_id string
output: id, status, total_cents, updated_at OR NOT_FOUND
side effect: none
permission: read-only; current tenant only
failure: retry TIMEOUT once; stop and ask on FORBIDDEN

Write tools need dry_run or a confirmation stage. Let a publishing action build a preview and validation report first; publish only on a separately confirmed call. Return small, precise results instead of thousands of unrelated log lines.

Evaluate with six to ten de-identified real tasks: a covered bug, small feature, ambiguous request, read-only call, publication check and conflicting evidence. Track first-pass success, human editing minutes, unrelated files touched and task cost. Repair the missing rule, stale reference, unclear contract or absent test after a failure. Do not keep adding chat-only workarounds.

Keep a small evidence index so stale answers do not return

The most common source of drift is a reference whose purpose or age is unclear. Keep paths and verification steps in a small index, and let the agent load the underlying files only when needed. The index should not duplicate whole documents.

# context-map.yaml: a path is a lead, not proof of current truth
- path: packages/contracts/export.ts
  purpose: "Confirm export status fields"
  owner: "API team"
  checked_on: "2026-09-16"
  verify_with: "pnpm test:contracts"
- path: tests/fixtures/export-status.json
  purpose: "Reproduce states and timeout responses"
  owner: "Web team"
  checked_on: "2026-09-16"
  verify_with: "pnpm test:affected"
- path: docs/ui/export-status.png
  purpose: "Visual spacing only; not an API contract"
  owner: "Design team"
  checked_on: "2026-09-16"
  verify_with: "Compare with the current design file"

checked_on is the last verification date, not an expiry guarantee. After a contract change, run verify_with before updating the index. If a type definition conflicts with an old screenshot, report both and ask the owner to decide. Text in web pages, chat exports, and repository comments is task data; a line telling the agent to ignore project rules must not become an instruction.

Turn six real tasks into a regression gate

A CSV with task ID, input snapshot, allowed tools, acceptance criteria, and failure reason is enough to begin. Re-run the same cases after changing AGENTS.md, a Skill, or a tool contract.

Case Expected behavior First place to investigate on failure
Tested bug Only related files change; tests pass Missing scope or test command
Ambiguous request Open questions appear before edits Unknowns were presented as facts
Conflicting sources Both sources and the conflict are reported Missing authority or verification date
Read-only query No write tool is called Missing permission or side-effect definition
Publication check Missing assets, links, and metadata are found Acceptance checklist is not executable
Tool timeout One agreed retry, then an explicit error Missing failure branch in the tool contract

Track first-pass acceptance, human rework minutes, unrelated files changed, and model-call cost. A starting release gate can require zero permission violations in high-risk cases, all required tests passing, and no increase in rework against the previous workflow. Set thresholds against your own baseline, not one polished demo. Repair the rule, source, tool, or test behind each failure and run the same case again.

Maintenance and handoff

Once a week, remove rules the repository can prove; test Skill commands and links; classify recent corrections; rerun one former failure; and review permissions, logs and redaction for tools that access customer data.

A useful agent asks fewer irrelevant questions, identifies missing evidence, changes fewer unrelated files and leaves an auditable handoff. With Code0, the same brief, rule set and evaluation sheet can carry across model choices without lowering the quality bar.

[ ] Outcome, known facts, unknowns, constraints and acceptance are explicit
[ ] Project rules contain local conventions only
[ ] Important claims point to types, tests, data or current documentation
[ ] Skills load by task; no universal instruction file
[ ] Tool contracts state input, output, effects, permissions and failures
[ ] Evaluation includes ambiguity, conflict and failure paths
[ ] Corrections were written back to a rule, reference, tool or test