TrustAI
Models Demo Research About us Privacy & Terms
Book a demo ↗ Sign in

Models

Managed models Demo

Company

Research About us Privacy & Terms Book a demo ↗

Featured

The small model already did the reading.

August 20, 2026 Qwen3 4B · Minitron 4B → Llama 3.1 8B

Before a model answers, it pays to read. We are making that prefill step cheaper by letting a smaller model read first, then carrying the work into Qwen and Llama targets.

Read blog
LLAMA 8K TARGET START · SMALL-MODEL CACHE READY 303 MS 38 MS LARGE MODEL READS WORK MOVES FORWARD 7.91× FASTER 7.91× faster Llama target start at 8K
All research notes

A learned KV handoff cut 8K target time-to-first-token by 87%.

August 18, 2026 Research note 007

A frozen Minitron 4B to Llama 3.1 8B study made an 8K target handoff 7.9× faster than native prefill while improving held-out cache quality over direct reuse.

Read blog
TARGET TTFT BY PREFIX LENGTH 1K 8K PREFIX NATIVE 303 MS HANDOFF 38 MS 7.9× 8K target TTFT speedup

A rollout residual cut 30B sibling-transfer divergence by 6.5%.

August 17, 2026 Research note 006

A frozen Qwen3-30B-A3B sibling test improved 20 of 28 paired windows, while a matching-layer ridge map produced the lowest mean divergence.

Read blog
28 PAIRED EVALUATION WINDOWS LOWER KL32 20 IMPROVED 8 OTHER 20/28 paired windows improved

Alignment first cut 32-step KV-transfer divergence by 11.1%.

August 16, 2026 Research note 005

Nine overnight Qwen3 sibling-transfer experiments showed that ridge alignment followed by rollout-aware correction produced our strongest 4B result.

Read blog
ALIGN STATE, THEN CORRECT BEHAVIOR .0670 .0429 .0381 DIRECT RIDGE + RESIDUAL 11.1% KL32 reduction after ridge

Partial recomputation reached 100% next-token agreement.

August 15, 2026 Research note 004

Recomputing two target layers removed the nonlinear ceiling of an affine cache map across 480 held-out toy-model comparisons.

Read blog
RECOMPUTE FROM THE EXACT BOUNDARY L3 · EXACT K/V PROJECTION L2 · TARGET RECOMPUTE L1 · TARGET RECOMPUTE L0 · BYTE-IDENTICAL 2 TARGET BLOCKS APPROXIMATE → FAITHFUL 100% next-token agreement

The mapped cache retained 91.2% of task quality.

August 10, 2026 Research note 003

Our Qwen3 1.7B to 4B map retained 91.2% of chance-normalized task quality and reached 65% next-token agreement.

Read blog
HELD-OUT TASK ACCURACY HELLASWAG PIQA ORACLE DIRECT MAPPED 65% next-token agreement

A ten-point sweep revealed three mapper tradeoffs.

August 10, 2026 Research note 002

A ten-point sweep picked k=24 on prefix loss while attention similarity peaked at k=8 and mapping time kept climbing.

Read blog
TEN-POINT MAPPER SWEEP PREFIX LOSS ↓ ATTENTION COSINE ↑ MAPPING TIME ↓ k=8 k=24 2.6× mapping time, k=8 to k=24

Moving a KV cache from Qwen3 0.6B to 1.7B.

August 9, 2026 Research note 001

A learned map retained 85.1% of chance-normalized task quality and improved next-token agreement 30× over direct injection.

Read blog
NEXT-TOKEN AGREEMENT 0.6B 1.7B KV MAP DIRECT INJECTION 2% LEARNED MAP 61% 30× OVER DIRECT 85.1% task quality retained
TrustAI

Novel inference, cheaper and faster.

Book a demo ↗

Products

Inference Demo Open-weight models

Company

About us Privacy & Terms
© 2026 TrustAI
LinkedIn X
Backed by Y Combinator Combinator