Tokenizers Stabilize, Skills Scale, and OpenAI Hits Bedrock

Tokenizers Stabilize, Skills Scale, and OpenAI Hits Bedrock

Oct 9, 2026 • 2 min read

The shift to GPT-6 is starting to reshape tooling assumptions, from tokenizers to cost models. This week brought updates that make agent development faster and more portable, including new runtime-level skill management and cross-cloud deployment options for OpenAI models.

Release ttok 1.0

Simon Willison bumped ttok to 1.0 and switched its default tokenizer to GPT-5/GPT-6, confirming that all seven tested GPT-5.5, 5.6, and 6 variants share identical tokenization at 44,794 tokens. This matters for anyone building token estimation into CI pipelines or cost forecasting tools: you can now assume GPT-6 uses the same token boundaries as GPT-5, which simplifies cost projection logic when you're evaluating model upgrades.

// Cost estimation now stable across GPT-5/6 family
import { countTokens } from 'ttok';
 
const prompt = fs.readFileSync('agent-prompt.txt', 'utf8');
const tokens = countTokens(prompt, 'gpt-6'); // uses same tokenizer as gpt-5
const estimatedCost = (tokens / 1_000_000) * inputPricePerMillion;

Release llm-openai-decisions 0.1a0

OpenAI's new Decisions API charges 10 cents per million input tokens with zero output cost, which is more than double Jev's 4.2 cents but eliminates output billing entirely. The llm-openai-decisions plugin supports yes/no, choice, and scoring question types with image inputs, making it useful for classification steps in agent workflows where you want predictable per-call costs. GPT-6 Astra built the plugin directly from API docs, which is a decent signal that the API surface is agent-friendly.

# Zero-output-cost classification in agent workflows
llm decisions yes-no "Does this PR modify cost-tracking logic?" \
  --input pr-diff.txt --model gpt-6-astra

Revamping Skills in Deep Agents

LangChain's skill system now supports lazy-loading tools, runtime pinning, and mid-thread reloads, addressing the context bloat problem as enterprise skill registries scale into the thousands. Binding tools to skills means only relevant context gets loaded, which keeps prompt caches valid and reduces token waste when agents cycle through specialized tasks. The ability to reload skills mid-thread without restarting is particularly useful for agent harnesses that need to pick up updated business logic during long-running workflows.

AWS Weekly Roundup: Amazon Bedrock Managed Agents powered by OpenAI

Amazon Bedrock now supports OpenAI models (GPT-6.1 Sol, GPT-6 Astra UltraFast, plus Claude Sonnet 5.5 and Grok 4.7) in managed agent infrastructure, which lets teams run OpenAI-based agents without routing traffic outside AWS. This matters for compliance-heavy environments where data residency or vendor lock-in prevention drives architecture decisions. Running OpenAI models through Bedrock also unifies billing and IAM with the rest of your AWS stack, which simplifies cost attribution when you're tracking agent spend across multiple model providers.

These updates collectively make agent development more predictable: tokenization is stable across GPT-5/6, new APIs offer fixed-cost decision primitives, skill systems scale without context explosions, and cross-cloud portability is improving. 🔧