Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics
AWS Samples releases context engineering techniques to prevent AI agent failures in long-running tasks
00

AWS Samples releases context engineering techniques to prevent AI agent failures in long-running tasks

Sep 13, 2026

The AWS Samples design guide for autonomous cloud coding agents highlights that shallow agents suffer from context overflow and goal loss during long-running tasks. To prevent these failures, frameworks like LangChain Deep Agents, Claude Code, and Amazon Bedrock AgentCore implement context engineering techniques within the agent harness. These mechanisms include token offloading, structured compaction, and todo-state recitation to manage finite attention budgets and maintain task objectives.

Agent harness context management

  • ▪AI agent frameworks including LangChain Deep Agents, Claude Code, Manus, OpenAI Codex, and Amazon Bedrock AgentCore implement context engineering mechanisms in their harnesses to sustain long-running tasks.
  • ▪The AWS Samples design guide for autonomous cloud coding agents states that shallow agents suffer from context overflow, goal loss, and failure to maintain state over long periods.
  • ▪The AWS Samples design guide states that the agent harness, which manages everything except the model, is the layer that prevents context overflow and goal loss.

Context window limitations

  • ▪Anthropic's context engineering guide states that attention creates n² pairwise relationships for n tokens, meaning every added token depletes a finite attention budget.
  • ▪The AI agent startup Manus reports that a typical task requires around 50 tool calls, causing original instructions to drift toward the middle of the context window where recall degrades.
  • ▪Chroma's Context Rot report evaluated 18 large language models and found that performance grows increasingly unreliable as input length grows, even on simple retrieval tasks.

Token offloading mechanisms

  • ▪LangChain Deep Agents offloads tool responses exceeding 20,000 tokens to the filesystem, replacing them with a file path and a 10-line preview.
  • ▪Claude Code returns any re-read file over 5,000 tokens as a path reference rather than full content after compaction.
  • ▪LangChain Deep Agents truncates older write and edit tool calls to a pointer when session context exceeds 85% of the model's window.

Context budgeting thresholds

  • ▪Claude Code caps auto memory at the first 200 lines or 25KB, and defers Model Context Protocol tool schemas by default, listing only tool names.
  • ▪Amazon Bedrock AgentCore uses a subagent architecture where a coordinator spawns three browser subagents in parallel in MicroVMs, and an analyst subagent receives only their structured findings.
  • ▪AWS reports that its AgentCore parallel subagent architecture has an expected runtime of 4 to 6 minutes, whereas sequential processing would take up to three times longer.

Compaction implementation strategies

  • ▪Following compaction, Claude Code re-reads up to five recently modified files, reloads matching rules, and re-injects invoked skill bodies capped at 5,000 tokens per skill.
  • ▪The Claude Developer Platform exposes a compact_20260112 context-management edit with custom instructions and a pause_after_compaction option.
  • ▪OpenAI's Responses API provides server-side compaction via context_management with a compact_threshold and a responses/compact endpoint that returns an encrypted compaction item.
  • ▪LangChain Deep Agents structures its compaction summary with dedicated fields for session intent, artifacts created, and next steps to preserve goals.
  • ▪Claude Code's compaction prompt discards redundant tool outputs while preserving architectural decisions, unresolved bugs, and implementation details.

Todo-state goal recitation

  • ▪An ETH Zurich study published in February 2026 found that repository context files like AGENTS.md do not improve task success while increasing inference costs by up to 23%.
  • ▪Manus agents maintain and rewrite a todo.md file step-by-step to recite objectives into the end of the context, reducing lost-in-the-middle drift.
  • ▪LangChain's Deep Agents v0.7 release in July 2026 made TodoListMiddleware opt-in after evaluations across three task categories showed slightly better reward and lower cost with todos disabled.

1 source

Marktechpost
Context Engineering Inside the Harness: 4 Mechanisms That Beat Context Overflow and Goal Loss on Long-Horizon Tasks
View source article

Featured stories

View more in Prompt engineering

OpenAI announces Codex cloud environments, Decisions API and Ultrafast tier at DevDay 2026

Sep 29, 2026 · 7 sources

OpenAI and Synopsys partner to develop AI model for chip design

Sep 30, 2026 · 3 sources

OpenAI launches Dots, always-on AI agents that work across 4,000+ apps

Sep 29, 2026 · 14 sources

OpenAI launches Space collaborative workspace and slides feature, competing with Microsoft office suite

Sep 29, 2026 · 13 sources

Story comments

Loading comments…

Related Projects

Amazon Web Services

Topics

Prompt engineeringAI for developersContext window managementAgent tool useAI agentsAI research & benchmarks

Featured stories

View more in Prompt engineering

OpenAI announces Codex cloud environments, Decisions API and Ultrafast tier at DevDay 2026

Sep 29, 2026 · 7 sources

OpenAI and Synopsys partner to develop AI model for chip design

Sep 30, 2026 · 3 sources

OpenAI launches Dots, always-on AI agents that work across 4,000+ apps

Sep 29, 2026 · 14 sources

OpenAI launches Space collaborative workspace and slides feature, competing with Microsoft office suite

Sep 29, 2026 · 13 sources