Articles / LLMs & Agents
Context Engineering for Agents: RAG, Memory, and Tool Results
Choose what an agent sees from retrieval, memory, and tool results with explicit freshness, provenance, and token budgets.

An agent fails long before it chooses a bad tool. It fails when its context mixes stale memory, irrelevant retrieval, unbounded tool output, and instructions from untrusted text. Context engineering is the practice of choosing, shaping, and discarding that information so the next model call has evidence it can use.
RAG, memory, and tool results are not interchangeable forms of context. They answer different questions:
| Source | Best for | Freshness | Main failure |
|---|---|---|---|
| RAG | Stable documents and policies | Index-dependent | Relevant text is not retrieved |
| Memory | User preferences and prior decisions | Must expire or be corrected | Old inference becomes a false fact |
| Tool result | Current state or exact computation | Usually request-time | Huge, private, or instruction-bearing output |
Anthropic's context-engineering article frames the objective as providing the smallest set of high-signal tokens for the next step. That is a practical constraint, not an aesthetic preference: excess context costs money, hides key evidence, and creates more material for a prompt injection to exploit.
Give every context item provenance and a budget
Represent context as records, rather than immediately concatenating strings. The code below is a complete, offline planning example. Token counts are rough word-based estimates for illustration; use the tokenizer of the model you deploy for an actual limit.
from dataclasses import dataclass
from datetime import datetime, timedelta
@dataclass(frozen=True)
class ContextItem:
source: str
text: str
observed_at: datetime
trusted: bool
token_estimate: int
def within_budget(items: list[ContextItem], budget: int) -> list[ContextItem]:
chosen: list[ContextItem] = []
spent = 0
for item in items:
if spent + item.token_estimate <= budget:
chosen.append(item)
spent += item.token_estimate
return chosen
def is_fresh(item: ContextItem, now: datetime) -> bool:
if item.source == "tool":
return now - item.observed_at <= timedelta(minutes=5)
if item.source == "memory":
return now - item.observed_at <= timedelta(days=30)
return True # RAG documents need their own version and validity metadata.
The trusted flag must not mean "safe to follow as an instruction." A retrieved support ticket and a third-party API response are data. Put them in clearly delimited evidence fields, tell the model they are untrusted, and never allow their text to redefine tool permissions or policy. Trust is also not binary in a real system: authorization, tenant, document version, and data classification may each need separate fields.
Use an order that reflects the next decision
A useful default assembly order is:
- Stable system policy and permitted tools, written by the application.
- The current user goal and authenticated scope.
- The freshest narrow tool result needed for the next decision.
- Retrieved passages selected for the current question, with source IDs.
- Compact memory that is relevant, attributable, and not expired.
This is not a prompt template. It is a selection policy. A current invoice total from a tool should beat a memory summary of the user's spending; a policy document versioned yesterday should beat a six-month-old chat summary. When sources conflict, present the conflict and ask a tool or person to resolve it. Do not blend them into a single authoritative paragraph.
Bound each source at the boundary
RAG should retrieve a small candidate set, filter by access and metadata, rerank if justified by evaluation, and retain document IDs with excerpts. Memory should store concise facts such as preferred_language=es with source and expiry, not a transcript of every conversation. Tool adapters should project only the fields required for the current action and cap rows, pages, and bytes.
For example, an order lookup can return order_id, status, updated_at, and the one field the user asked about. It should not return a 2,000-row audit log to let the model decide what matters. If follow-up detail is needed, expose a second bounded read tool keyed by order_id. This is the same capability discipline used in Build a Safe Read-Only MCP Server in Python.
Plan a fixed budget for a tool-using turn
Suppose an agent has a 6,000-token input budget after system instructions. Reserve it before calling components:
| Slice | Budget | Policy |
|---|---|---|
| Current task and state | 800 | Always include |
| Latest tool result | 1,200 | Project and truncate server-side |
| Retrieval evidence | 2,500 | Top passages with citations |
| Relevant memory | 600 | Only fresh, attributable facts |
| Working headroom | 900 | Leave room for the next observation |
These numbers are an example, not a universal ratio. Log which items were selected, excluded, and truncated by source type. Then evaluate tasks where the answer was wrong or the agent looped. If correct evidence was excluded, adjust retrieval or allocation. If too much evidence was included, improve selection before raising a context-window limit.
Keep the loop bounded and observable
After every tool call, summarize or store structured observations outside the model transcript, then feed only the piece relevant to the next action. Limit tool rounds and stop with a useful partial result or escalation when the budget expires. An agent should not repeatedly retrieve the same documents or reread an unchanged tool response because it has no state model.
RAG evaluation provides evidence for one slice of this system. Offline RAG evaluation separates missed retrieval from unsupported claims. For the overall agent, evaluate the final outcome, tool choices, information actually selected, and safe behavior under stale or adversarial context. Context design improves the odds of a useful next action; it does not turn untrusted text into authority or guarantee correctness.
About the author
Rodrigo Arenas is a software architect and machine learning engineer in Medellín, Colombia. He builds products, platforms, and open-source software for AI, Data Engineering, and Machine Learning: creator of Ciaren, sklearn-genetic-opt, and PyWorkforce.
Are you building an AI or data platform?
I design and build critical systems end to end: architecture, data, models, and product. If that is the scale you are working at, let us talk.