Skip to main content

Strands Harness: Test Its Agent Defaults and 28% Token-Cost Claim

4 min read

Strands released an open-source agent harness with context, tools and sessions. Here is how to test its company-reported 28% token-cost saving.

Strands Harness: Test Its Agent Defaults and 28% Token-Cost Claim

Building an agent is easy to demonstrate and harder to operate. Strands has released a ready-made harness for the parts around the model: tools, context, memory, and a repeatable execution loop. Its efficiency claim is worth testing, not repeating as a guarantee.

Strands launched its harness on September 21 as an Apache-2.0, open-source agent runtime available in Python and TypeScript. The project provides preconfigured defaults for tool use, context management, sessions, and memory, and can run locally or be deployed to cloud infrastructure. Its launch report says it supports models through Amazon Bedrock, Anthropic, OpenAI, Google, Ollama, and LiteLLM. Strands reports 28% lower token cost across six benchmarks using the same models while retaining nearly equal scores. Those are the project’s measurements; we didn’t see an independent reproduction in the material we reviewed.

A harness is more than a model wrapper.

A model call produces text or a tool request. A harness decides how that request is executed, what result comes back into context, when conversation history is compacted, and how a long-running task resumes. Those defaults can change both reliability and cost even when the underlying model is identical.

Strands says its default context policy truncates tool results above roughly 1,500 tokens, triggers summarization when the context window is more than 85% full, and attempts recovery inside the loop after an overflow. These choices reduce repeated context, but they can also hide a detail an agent needed later. The right question is whether the retained information still supports the task, not simply whether the token bill falls.

How to test the 28% claim

Strands’ reported comparison covers six benchmarks. The company says the harness used the same Claude or GPT models in its cost comparison and produced nearly equal scores, but the result is not a universal saving for every workflow. It also says it plans a follow-up research paper. Until that methodology is fully available, a buyer should run a small, matched evaluation on its own tasks.

MeasureWhy it matters
Model and versionA model change can overwhelm the effect of harness defaults.
Task set and success rubricLower cost has little value if completed work gets worse.
Tool access and permissionsDifferent tools can make one agent appear more capable.
Input, output and cached tokensA single total hides where savings occur.
Retries and human correctionsThese are real costs even when the API bill is lower.
Keep these variables fixed when comparing harnesses.

Run each task several times because tool choices vary. Save traces and compare the exact point where context is trimmed or summarized. If the cheaper run omits a requirement and someone must fix it, count the repair. This approach turns a benchmark headline into a deployment decision.

Permissions still live outside the benchmark.

The SDK repository and quickstart make it relatively easy to connect a model and tools. Production use needs a separate decision about what each tool may read or change. A harness can manage the execution loop without supplying your organization’s approval policy for file writes, credentials, or external systems.

Start with a constrained task with an objective result, a fixed tool set, and a reversible workspace. Then add session persistence, memory, and remote tools one at a time. This isolates whether a failure came from model reasoning, a tool boundary, or context management. Our guide to tool-level access explains why granting an agent a connection is not the same as authorizing every object behind it.

Who should try it

Strands is most relevant to teams building their own general-purpose agent that already know what tasks, tools, and approval steps they need. It provides a useful starting loop and concrete context defaults. It is not evidence that one line of setup makes an agent production-safe or that the reported savings will transfer to every model and task.

Sources

Leave a comment

Your email address will not be published. Required fields are marked *