The conversation around production agent systems is coalescing around a few hard problems: how to route intelligently to cut costs, how to combine specialized decision models with orchestration, and how to prevent infrastructure chaos as agent fleets grow. These four posts show where the tooling is headed.
Muse is getting a lot of attention...
Meta's Muse gives every consumer user a persistent Linux VM and an agentic interface that can act autonomously across it. The accessibility is groundbreaking, but John Gruber's power saw analogy hits the mark: people understand physical tools can hurt them, but may not grasp what an agent with VM access can actually do. This is the canary for a broader question the industry hasn't answered: what does informed consent look like when you hand someone an autonomous system?
How to Build a Model Router in the Harness
LangChain's Open SWE agent cut median task cost by 64% using a three-tier routing strategy: 56% of tasks go to a balanced tier, 34% to fast models, and only 10% need frontier performance. The key insight is that routing belongs in the harness, not a dumb gateway, because effective decisions require task and domain context. If you're running agents in CI or production workflows, this is the pattern to copy.
// Router in harness with task context
function selectModel(task: Task, history: ExecutionHistory) {
if (task.complexity === "high" || history.failedAttempts > 0) {
return "gpt-4o"; // performance tier
} else if (task.estimatedTokens > 5000) {
return "claude-3.5-sonnet"; // balanced tier
}
return "gpt-4o-mini"; // fast tier, 34% of traffic
}Building Prod with Jev and LangGraph
Jev is a decision model that handles structured choices 200x faster and 400x cheaper than general LLMs by returning typed answers instead of generating text. LangGraph provides the orchestration layer to wire Jev's fast decisions into durable workflows with state management and human-in-the-loop gates. This is the "unbundling of intelligence" in practice: route narrow decisions to cheap specialized models by default, escalate to frontier LLMs only when complexity demands it.
How to scale agentic applications without creating AI sprawl
Databricks frames the enterprise agent scaling problem as infrastructure sprawl: every new agent duplicates integrations, governance, and observability. Their answer is a shared foundation (Agent Bricks) providing model choice, governed data access, and centralized tracing across agent fleets. The pitch is less about their specific product and more about a necessary pattern: if you're deploying more than a handful of agents, you need shared infra or you'll rebuild the same pipes forever.
These posts sketch the same production pattern from different angles: specialize models to cut cost, build routing and orchestration into the harness where context lives, and centralize infrastructure to avoid sprawl. The focus is shifting from "can we build agents" to "can we run them economically at scale." 🔧