Home
Blog
AI Agent Context Explained: What Agents Can't See in Your Infrastructure

AI Agent Context Explained: What Agents Can't See in Your Infrastructure

Zeen Rachidi
Product Marketing
with special guest
Mitchell
Hashimoto
Mitchell Hashimoto headshot

The "C" word is a controversial subject in the US, but we have to talk about "context", and what it means to an AI Agent. For starters, agents can only act on what's in their context window. Everything outside of it is a guess. In application code that limit is usually an annoyance. In infrastructure it's expensive, because the things an agent can't see (like resources created in the console, dependencies no repo mentions, or deployments nobody owns anymore) are the things that cost money and cause headlines.

This article explains what agent context is, how it works, and where it stops working. It then applies those limits to a problem platform teams are already facing: cloud resources that nobody owns.

What "context" means for an AI agent

An AI agent is a language model running in a loop. On each turn it receives a prompt and produces the next step: text, a tool call, or a finished answer. For Claude, that prompt is assembled in a fixed order: tool definitions, then system instructions, then the conversation messages. Tool results land in the messages, so everything the agent has read so far (files, command output, API responses) sits in that one prompt.

The model reads the prompt as tokens, small units of text, and predicts what comes next. The context window is the maximum number of tokens it can take in at once. Anything outside the window does not exist for the model on that turn.

The conversation is also re-sent. In Claude Code, for example, resuming a session sends the whole conversation again, and caching lets the API skip reprocessing the part that has not changed.

How attention works, and why it limits context

Modern language models are built on the transformer architecture. Its key mechanism, attention, lets each token weigh the other tokens in the prompt when the model works out what it means. In decoder-only models, the design of the open models Liu et al. evaluated, a token can look only at the tokens before it.

That has two practical consequences, shown in the diagram below.

Longer prompts cost disproportionately more. Each token is compared with itself and every token before it. With 4 tokens that is 10 comparisons (1 + 2 + 3 + 4). With 8 tokens it is 36. Doubling the prompt nearly quadruples the work, which the original transformer paper expresses as O(n²·d) per layer, where n is the sequence length.

Faster code changes the constant, not the shape. FlashAttention is a smarter way to run the same calculation. It returns exactly the same result while moving far less data in and out of GPU memory, and its memory use grows linearly with length. The number of comparisons does not shrink, though, and a recent survey notes that the compute stays quadratic.

Attention cost grows with the square of prompt length. Optimizations such as FlashAttention reduce memory use and memory traffic, not the number of comparisons.
Attention cost grows with the square of prompt length. Optimizations such as FlashAttention reduce memory use and memory traffic, not the number of comparisons.

Caching solves a different problem. When the same prompt prefix is sent again, caching lets the provider skip reprocessing it. On Anthropic's API, cache reads cost a fraction of the base input price (0.1x on most models, lower on a few), while writes cost 1.25x for a 5-minute cache and 2x for a 1-hour cache. The same documentation states that caching has no effect on output generation: the response is identical to what an uncached request would return. Caching makes context cheaper and faster. It does not make the model notice more of it.

Five limits that matter for agents

1. The window is a hard boundary

An agent can check a change against what it can see. It cannot check against what it cannot. Where the missing fact matters, the agent has to infer, and inference is where confident mistakes come from.

Here is a hypothetical, offered as an illustration and not as a documented incident. A platform team asks an agent to remove unused security group rules from a Terraform repository. The agent finds this rule, searches the repository for anything that references it, finds nothing, and deletes it:

resource "aws_security_group_rule" "reports_to_db" {
type = "ingress"
security_group_id = aws_security_group.shared_db.id
protocol = "tcp"
from_port = 5432
to_port = 5432
cidr_blocks = ["10.20.0.0/16"]
}

The rule was the only path for a reporting service that another team created in the console. The repository cannot tell the agent that, because the dependency exists in the cloud and not in the code. No amount of careful reasoning inside the window recovers a fact that was never in it.

A dependency that exists only in the cloud is invisible to an agent working from the repository.
A dependency that exists only in the cloud is invisible to an agent working from the repository.

2. A bigger window is not a fix

Cost is the first problem, as above. Quality is the second. Liu et al. found that model performance can degrade significantly depending on where the relevant information sits in the prompt, often lowest when it is in the middle, including for models built for long contexts. Chroma's 2025 report found consistent degradation as input length grew across the models it tested, after confirming the models handled the focused version of the same task.

Two caveats. Liu et al. tested older models, including GPT-3.5-Turbo and Claude-1.3. Chroma's report is industry research, not peer reviewed. Newer models may behave differently, so check current evaluations before relying on any long-context claim. The defensible takeaway is narrower: length is not free, and more context can mean worse use of it.

Schematic sketches of the two effects. Shapes are illustrative, not measured data.
Schematic sketches of the two effects. Shapes are illustrative, not measured data.

3. Cheaper attention trades away links

The transformer paper itself describes restricting attention to a neighborhood of size r. Per-layer cost drops to O(r·n·d), but the maximum path length between positions rises to O(n/r). In plain terms, connecting distant parts of a prompt takes more steps. That is a trade-off, not a free saving, and in a large system the distant connections are often the ones that matter.

Restricting attention lowers cost but lengthens the path between distant tokens (O(n/r)).
Restricting attention lowers cost but lengthens the path between distant tokens (O(n/r)).

4. Retrieval finds what's written down

Many agents work around the window by searching a codebase or index and pulling relevant pieces into context. That works when the connection is written in something searchable. A dependency that lives in runtime state, a console-created resource, a shared database reached by hostname, or an ownership agreement that lives in someone's head leaves nothing to search for.

5. Context is a metered resource

Every token is billed, and caching depends on an exact prefix match. Anthropic's guidance on reducing cost notes that a dynamic timestamp in the system prompt can break the cache. So context hygiene is also cost control: stable instructions first, volatile content last, and only what the task needs.

A change early in the prompt invalidates the cached prefix from that point on. Placing changing values last preserves cache hits.
A change early in the prompt invalidates the cached prefix from that point on. Placing changing values last preserves cache hits.

Where this meets infrastructure

Code describes what a team intended. The cloud holds what exists. env zero's EZ Control announcement describes everything outside the infrastructure as code workflow (ClickOps, break-glass fixes, automated pipelines, and increasingly AI agents) as creating resources that IaC tools never see. An agent that reads the repository sees the first set and not the second.

Ownership sits in that gap. Who created a resource, which team pays for it, and what depends on it are facts about runtime state and organizational agreements. The Cloud Accountability Model sets the standard plainly: every resource needs an accountable owner, not just a creator. Why Governance Fails in Multi-Team Cloud Environments lists poor visibility into what exists, who owns it, and what it costs among the common failure points.

What do the numbers say? Flexera's 2026 State of the Cloud Report surveyed 753 cloud decision-makers and puts estimated wasted IaaS and PaaS spend at 29%, reversing a five-year decline. Flexera attributes the uptick to cost complexity from AI and new services. Its five-year view puts the estimate at 30% in 2021 and 29% in 2026. In the same report, FinOps team adoption reached 63% and CCoE adoption 71%.

We have to keep in mind that these are respondents' estimates, not measured spend, and Flexera does not isolate forgotten deployments as a cause. What the data does show is that governance structures have grown while estimated waste has not fallen. Two conclusions could be drawn from this: agents can create resources faster than a person can tag them by hand, and an agent that can't see ownership can't avoid touching something ownerless.

What to do about it

Put ownership on the resource, at creation

Ownership that lives in a wiki is invisible to an agent. Ownership that lives on the resource can be read, queried, and enforced. In the AWS provider, default_tags applies tags to supported resources from one place (Auto Scaling Groups are the documented exception):

provider "aws" {
region = "us-east-2"

default_tags {
tags = {
owner = "team-payments"
managed_by = "terraform"
expires = "2026-12-31"
}
}
}

Then enforce it. A policy that blocks a change when the owner tag is missing works the same for a person and for an agent, which is the property AI agent governance depends on.

Give agents live state, not only code

An agent working from code alone sees intent. An agent working from structured, current state sees what exists. env zero's Agentic Experience launch makes the same point: without a real read surface, an agent sees a thin slice of state, fills gaps with guesses, and spends tokens scraping an API that was never built for a machine to reason over. Its env0 context command returns an environment's state, recent deployments, and drift or failure summary in a single call.

Curate the window

Send the smallest useful set. Give ownership and dependency facts as compact structured data instead of pasting whole files. Keep stable instructions at the start and volatile values out of the cached prefix.

Verify outside the model

A model cannot verify what it cannot see, so the checks that matter sit outside it: plan review, policy as code, approval gates, and drift detection. The AI agent governance guide covers how to build these so they hold no matter who or what made the request.

Expire what agents create

That guide also recommends TTLs on anything an agent creates. Expiry works best when ownership is defined: in the Cloud Accountability Model, Salt Security's scheduled weekend shutdowns of non-production environments were possible because each environment had a defined owning team.

Rediscover on a schedule

Inventory decays. The Terraform bulk import guide makes the same point for unmanaged resources: treat discovery as recurring, because new console-created resources accumulate the way the original backlog did.

How env zero approaches this

EZ Control, now in Early Access, discovers cloud and SaaS resources continuously, reconciles them against what code declared, and structures the result into a context layer. Each resource is linked to the code that declared it, the team that owns it, its cost, its dependencies, and the policies that apply. The initial connection is agentless and read-only, so teams can start by observing before granting any authority to act.

Through the Agentic Experience, agents read real environment state under scoped identities, with the same approval gates a person would face. Cloud Compass categorizes resources by how they were actually changed (console, API or CLI, or IaC) and scores each by severity.

Want to see how env zero gives agents and platform teams the same picture of your cloud? Schedule a demo.

Key points

  • An agent acts only on what is in its context window. Missing facts are inferred, not looked up.
  • Bigger windows cost more, and published evaluations show quality can fall as input grows. Caching cuts cost and latency, not what the model can use.
  • Code describes intent. Ownership, dependencies, and live state usually live outside it.
  • Put owner and expiry on the resource, give agents structured live state, and verify outside the model.

Frequently asked questions

What is an AI agent's context window?

It is the maximum number of tokens the model can take in for one request. It includes tool definitions, system instructions, and the whole conversation, including tool results. Anything outside it is invisible to the model on that turn.

Does a larger context window solve the problem?

Not on its own. Attention cost grows roughly with the square of input length, and published evaluations (Liu et al., Chroma) found performance can degrade as input grows or as relevant information moves to the middle. Newer models may differ, so check current evaluations.

Does prompt caching fix context limits?

No. Per Anthropic's documentation, caching reuses a stored prompt prefix to cut cost and latency and does not change the model's output. It does not expand what the model can see or improve how well it uses it.

Why does resource ownership matter for AI agents?

Ownership and dependency facts usually live in cloud runtime state or in people's heads, not in code. An agent working from a repository cannot see them, so it can change or create resources without knowing who owns them or what depends on them.

What should an agent be able to see about my infrastructure?

At minimum, the current state of the resources it is about to change, who owns them, what depends on them, and which policies apply. Structured, queryable state is better than raw API output, and read-only access is a safe starting point.

Schedule a technical demo
See env zero in action
Schedule demo

Related Content

All articles
Read more
Read more
Read more
Read more
Read more
Read more