@takk/racs - v1.0.0 - Apache-2.0

Cache the prefix.
Bank the difference.

RACS (Remote Agent Context Store) plans provider-faithful prefix-cache directives for Massive Intelligence (IM) agents without ever calling a provider API. Declare the prompt as segments; RACS lints the layout for the documented cache-killers, emits the exact directive your provider understands, schedules keep-warm refreshes, detects prefix drift, and reports savings in your own pricing. Zero credentials, zero network, the host stays in control.

pnpm add @takk/racs
187tests passing
92%coverage
0runtime deps
16provider profiles
The engine

Five components and one invariant.

Prefix caching fails in production for boring reasons: a timestamp in the system prompt, a tool list that reorders itself, a layout that puts the volatile part first. An OpenClaw issue measured ten times the expected cost from exactly that failure mode. RACS makes each of these a planned, linted, measured decision.

Stability analyzer

Nine lint codes catch the documented cache-killers before they bill you: timestamps and identifiers in segments declared stable, volatile-early layouts, unstable tool definitions, below-minimum prefixes, write-premium traps. Errors defeat caching, warnings cost money, and racs analyze exits nonzero in CI.

The silent cache-bust class of incident becomes a failed build instead of a surprise on the invoice.

Directive planner

One plan() call resolves the provider profile and emits the directive that provider actually understands, across four adapter families: breakpoint (cache_control, cachePoint), routing-key (prompt_cache_key), resource (cachedContents), and passive (ordering advice only). Deterministic FNV-1a 64 prefix keys and break-even math come with every plan.

16 providers, 4 code paths, zero per-provider conditionals in your application.

Usage ledger

Feed it the usage counts your provider already returns and it normalizes them into hit ratios, read and write token totals, and USD savings, computed strictly from pricing you supply. RACS never hardcodes a provider price.

One stats() call answers the question finance actually asks: what did caching save this month, net of write premiums.

Drift detection

Every plan fingerprints its stable segments. When the same agent lineage produces a different prefix key, the drift report names exactly which segments changed and how many cached tokens died with the change.

A one-byte edit to a "stable" prompt is caught on the next plan, not three invoices later.

Refresh scheduler

Provider caches expire on a TTL; a touch shortly before expiry keeps them warm at read price instead of paying the write premium again. schedule() computes the due moments at 90 percent of each window; the host runs the timer and the call.

Keep-warm stops being a hand-rolled heartbeat script and becomes a library primitive.

The invariant

RACS never calls a provider API. It holds zero credentials, opens zero network connections, and never sends a byte anywhere. Planning and accounting are synchronous and in-process; the host applies every directive and makes every call itself.

Nothing new to audit: no credential to store, no egress path to review, no vendor in the request loop.

Provider matrix

16 providers, 4 families, 1 table.

Providers disagree on cache semantics: Anthropic wants explicit breakpoints, OpenAI caches automatically behind routing keys with no write counter, Gemini sells the cache as a server-side resource with storage economics. RACS models each provider as a thin profile over one of four adapter families, and every profile value is overridable per engine instance.

breakpoint

Explicit breakpoints, write premiums

The caller marks cache boundaries inside the request body, up to 4 per request, with 5m and 1h TTL tiers. Writes cost 1.25x base input on the 5m tier and 2x on the 1h tier; cached reads cost 0.1x. Plans place the breakpoints and prove they pay back.

routing-key

Automatic caching, steered routing

The provider caches server-side; the caller can only steer routing with a key so identical prefixes land on the same cache. No write premium, reads around 0.25x, optional 24h retention. Plans emit the sticky key and the layout that makes it hit.

resource

Cache as a server resource

The cache is created, refreshed, and deleted by the host, with caller-set TTL and per-token-hour storage billing. Plans emit the full lifecycle: create, read, refresh, delete, plus the storage economics that decide whether the resource is worth keeping.

passive

No control surface, still optimizable

The provider caches automatically and exposes no controls. RACS still orders segments stable-first, lints the layout, and accounts usage, because on a left-anchored cache the ordering itself is the optimization.

breakpoint anthropic

cache_control, 4 breakpoints, 5m and 1h TTLs, 0.1x reads

breakpoint bedrock

cachePoint on the Converse API, Anthropic-equivalent semantics

breakpoint hermes

Hermes Agent system_and_3 layout over cache_control semantics

breakpoint microsoft-foundry

Claude on Microsoft Foundry honors cache_control unchanged

routing-key openai

prompt_cache_key, 128-token increments, optional 24h retention

routing-key xai

x-grok-conv-id and prompt_cache_key, reads around 0.16x

routing-key mistral

64-token cache blocks, prompt_cache_key routing, 0.1x reads

routing-key moonshot

Kimi platform caching through the OpenAI-compatible surface

routing-key openrouter

cache_control passthrough and cached_tokens across upstreams

resource google

cachedContents lifecycle, caller-set TTL, storage per token-hour

passive groq

automatic on gpt-oss models, entries expire after 2 hours idle

passive deepseek

disk-based context cache, hit and miss token reporting, 0.1x reads

passive ollama

local runtime KV reuse, analytics measure latency, not billing

passive lmstudio

local runtime KV reuse, same posture as ollama

passive huggingface

no public prefix-cache controls as of June 2026, ordering still plans

passive custom

extensible: define your own profile via options.profiles, any family

Quickstart

Plan, apply, record, measure.

1. Add the package

pnpm add @takk/racs
npm install @takk/racs
yarn add @takk/racs
bun add @takk/racs

Zero runtime dependencies. No peer dependencies. The tarball is 46 files; the core entry is 10.72 kB brotli.

2. Plan the cache for one prompt

import { createRACS } from '@takk/racs';

const racs = createRACS({
  pricing: {
    // your prices, USD per million tokens; RACS never hardcodes them
    'claude-sonnet-4-5': { inputPerMTok: 3, cacheReadPerMTok: 0.3, cacheWrite5mPerMTok: 3.75 },
  },
});

const plan = racs.plan({
  agentId: 'support-agent',
  provider: 'anthropic',
  model: 'claude-sonnet-4-5',
  segments: [
    { id: 'system', role: 'system', stability: 'stable', content: SYSTEM_PROMPT },
    { id: 'tools', role: 'tools', stability: 'stable', content: TOOL_DEFINITIONS },
    { id: 'turn', role: 'dynamic', stability: 'volatile', content: userMessage },
  ],
  reuse: { intervalSeconds: 45 },
});

plan.directives; // provider-faithful: cache_control breakpoints, in order
plan.findings;   // lint findings; errors mean the layout defeats caching
plan.breakEven;  // write premium vs read savings, stated in tokens

3. Record usage, read the savings

// after the provider call, report the counts the response already carries
racs.record({
  provider: 'anthropic',
  model: 'claude-sonnet-4-5',
  prefixKey: plan.prefixKey,
  inputTokens: 5200,
  cacheReadTokens: 4100,
});

racs.stats(); // hit ratio, read/write tokens, USD saved, net of premiums

// keep-warm: entries come due at 90 percent of each TTL window
for (const entry of racs.schedule()) {
  // the host touches the cache (a cheap read), then reports it
  racs.markRefreshed(entry.prefixKey);
}
CLI

The proof runs in your terminal.

The racs binary ships in the box: help, version, analyze for CI gating, simulate for the deterministic demonstration below, inspect --watch for live state, and serve, a hardened local HTTP bridge exposing /plan, /lint, /usage, /stats, /schedule, /refreshed, and /invalidate to non-JavaScript stacks.

$ racs simulate --calls 400 --seed 7
racs simulate: 400 calls, seed 7, interval 60s, provider anthropic
structured lint: clean
naive lint:
LINT warning segment-order naive-turn Volatile segment 'naive-turn' precedes stable segment 'naive-tools'. Prefix caches are left-anchored, so every token after 'naive-turn' is unreachable for the cache. [...]
LINT warning timestamp-in-stable naive-system Segment 'naive-system' is declared stable but contains an ISO-8601 datetime (digest 9d6f8366), and the words 'today' or 'current time' near digits. [...]
drift naive: 217dbd595cc63a93 -> 7dfee9d9ed5102b9, segments [naive-system], 3000 tokens invalidated (call 2)
drift naive: 7dfee9d9ed5102b9 -> 44338bff92d176f9, segments [naive-system], 3000 tokens invalidated (call 3)
[... 400 lines hidden: 397 further drift reports (calls 4 through 400) and the progress lines at calls 100, 200, and 300 ...]
progress: 400/400 calls, structured hits 399, naive hits 0
--- summary ---
calls: 400
structured: hit ratio 0.96, net savings 9.87 USD
naive: hit ratio 0.00, write-premium loss 1.50 USD
structured prompt saves $11.37 (88.1%) versus naive

Output of racs simulate --calls 400 --seed 7 against the shipped demonstration pricing, truncated for length. Same seed, same output, every time: the simulation is fully deterministic.

Savings calculator

What would your stable prefix save?

Measured agent workloads report 41 to 80 percent savings on input spend from prefix caching (arXiv 2601.06007, January 2026). Your number depends on three quantities you already know. Estimate it from the provider multipliers.

Input tokens only; output spend is untouched by prefix caching.
60%
System prompt, tool definitions, reference documents: the part that repeats verbatim across calls.
70%
How often a call replays the prefix inside the TTL window. Dense agent loops run high; sporadic traffic runs low.

Estimated monthly savings

$333.00

33.3% of input spend

Cache-read savings
$378.00
Write premium paid on misses
$45.00

An estimate from the provider multipliers and your own numbers, nothing more. Server-side truth comes from usage reports: record real calls with record() and read stats().

Surfaces

Seven ways in, one engine underneath.

Sizes are brotli-compressed, measured on the published build and enforced in CI by size-limit.

@takk/racs

10.72 kB ESM / 10.86 kB CJS

The full engine: plan, lint, record, stats, schedule, markRefreshed, drifts, invalidate, profileOf, on, flush, close, plus the memory, file, and KV-style state backends.

@takk/racs/otel

604 B

GenAI span ingestion: maps OpenTelemetry GenAI semantic-convention spans into normalized usage records, so traces you already collect feed the ledger without new instrumentation.

@takk/racs/vercel

1.15 kB

Middleware for the Vercel AI SDK: applies plan directives through providerOptions and records usage from results, streaming included via wrapStream.

@takk/racs/integrations

666 B

Bridges to the four sibling packages: freeze parameter tuning while the prefix drifts and release after 3 stable plans, observe the cache behaviorally, plan per routed model, invalidate on credential rotation.

@takk/racs/web

10.23 kB

The same engine without the Node file backend, so browser bundles never advertise a backend they cannot run. Pair with kvState for persistence.

@takk/racs/edge

10.23 kB

Edge-runtime entry with the identical surface. kvState wraps any Redis, Upstash, or Cloudflare KV style client in one line for cross-isolate state.

CLI: racs

help / analyze / simulate / inspect / serve

analyze gates prompt changes in CI, inspect --watch redraws live state, and serve is a hardened, token-authenticated HTTP bridge for non-JavaScript stacks.

Engine internals

exported, not hidden

Planner, PrefixAnalyzer, Ledger, PROVIDER_PROFILES, fnv1a64, and estimateTokens are all public exports; build your own pipeline from the same parts the engine uses.

Telemetry

6 typed events

plan.created, prefix.drifted, usage.recorded, refresh.due, resource.action, limit.reached: synchronous, deterministic under an injected clock, and free of any telemetry vendor.

Family stack

Five packages, one operational layer for agents.

Each package stands alone with zero runtime dependencies; together they cover credentials, routing, behavior, tuning, and caching. The /integrations entry bridges them in one line each.

@takk/keymesh

Credential rotation, failover, and circuit breaking for provider API keys; RACS invalidates cached resources when a key rotates, via keymeshBridge.

davccavalcante/keymesh
@takk/modelchain

Model routing and fallback chains across providers; modelchainBridge plans a cache layout for whichever model the chain routes to.

davccavalcante/modelchain
@takk/behavioralai

Behavioral observation for agent systems; behavioralaiBridge turns the cache into a behaviorally observed surface, hit and drift patterns included.

davccavalcante/behavioralai
@takk/noeticos

Adaptive runtime parameter tuning with canary rollouts and rollback; noeticosBridge freezes tuning while the prefix drifts and releases it after 3 stable plans.

davccavalcante/noeticos
@takk/racs this package

Prefix-cache management: stability linting, provider-faithful directives, drift detection, keep-warm scheduling, savings analytics.

davccavalcante/racs
FAQ

Common questions.

Is RACS production-ready at 1.0.0?

Yes. 187 tests across 13 suites pass on Node 22 and Node 24, with coverage at 92.46 percent statements, 82.09 percent branches, 93.33 percent functions, and 92.49 percent lines. TypeScript strict mode, lint, and package checks are clean, and every release ships with SLSA provenance.

Does RACS ever call my provider's API?

Never. RACS plans directives, lints layouts, schedules refreshes, and accounts usage entirely in-process. It holds zero credentials and opens zero network connections; the host application stays in control of every provider call and applies the directives itself.

How is RACS different from a gateway or proxy?

A gateway sits on the wire and forwards traffic. RACS is a planning engine: it computes the cache directives your existing client applies on its own requests. As of June 2026 we found no shipping npm package that combines stability linting, multi-provider directive planning, drift detection, persistence, and savings analytics in one library.

Which providers does RACS support?

16 profiles over 4 adapter families. Breakpoint: anthropic, bedrock, hermes, microsoft-foundry. Routing-key: openai, xai, mistral, moonshot, openrouter. Resource: google. Passive: groq, deepseek, ollama, lmstudio, huggingface, custom. Every profile value is overridable per engine instance, so a provider changing terms is a configuration edit, not a release.

How much can prefix caching actually save?

Measured agent workloads report 41 to 80 percent savings on input spend (arXiv 2601.06007, January 2026). Your number depends on prefix share, hit ratio, and the provider multipliers; the calculator on this page estimates it from your own figures, and the ledger reports the measured truth once you record real usage.

What do the nine lint codes catch?

The documented production cache-killers: timestamps and identifiers inside segments declared stable, volatile content placed early in the prompt, unstable tool definitions, prefixes below the provider minimum, breakpoints placed after volatile segments, write-premium traps, wrong segment ordering, and missing stability declarations.

Does RACS see my prompt content?

Only if you pass it. Every segment accepts a contentHash instead of content. In hash-only mode RACS never sees the text, and plans, drift reports, persisted snapshots, and telemetry carry hashes and token counts only.

How does drift detection work?

Every plan fingerprints its stable segments with deterministic FNV-1a 64 hashing. When a later plan for the same agent lineage produces a different prefix key, RACS emits a drift report naming exactly which segments changed and how many cached tokens were invalidated.

Where does state live, and can it persist?

In memory by default. The file backend persists snapshots on Node, and kvState wraps any Redis, Upstash, or Cloudflare KV style client in one line. Persistence is optional; the engine is fully functional without it.

How are USD savings calculated?

From pricing you supply. RACS never hardcodes provider prices: you pass a table in USD per million tokens, and the ledger combines it with normalized usage reports into hit ratios, savings, and net effect after write premiums. Without a pricing table, every statistic is still reported in tokens.

Author

Built and maintained by David C Cavalcante.

David C Cavalcante

Founder, Takk Innovate Studio

Builder of the @takk family of NPM packages: zero-dependency operational infrastructure for Massive Intelligence (IM) agent systems.

RACS exists because the cheapest tokens are the ones the provider already has. The Hermes Agent ecosystem alone carries multiple open issues asking for stability linting, drift detection, and cache accounting; this package answers them as a library any host can embed.

If RACS catches a cache-buster the test suite missed, the most useful thing you can do is open a GitHub issue with the lint finding attached. The release runbook, the threat model, and the contributor agreement live in the repository.