Together AI Open-Sources OSCAR: An Attention-Aware 2-Bit KV Cache Quantization System for Long Context LLM Serving

Long-context inference makes the KV cache one of the main costs of serving LLMs. During autoregressive decoding, the cache grows with context length, batch size, and model depth. At high batch sizes and long contexts with 100K tokens across dozens of concurrent requests, the KV cache consumes a large fraction of GPU memory. Compressing it is a direct way to increase batch size and reduce memory traffic.

The obvious approach is quantization. However, pushing KV caches to INT2 (2-bit) precision has been largely impractical. Prior methods either collapse in accuracy or require custom serving layouts incompatible with paged KV-cache systems. Together AI’s OSCAR (Offline Spectral Covariance-Aware Rotation) addresses both problems.

Why INT2 KV Cache Quantization is Hard

KV activations contain channel-wise outliers. A small subset of channels holds extremely large values, while most channels are well-behaved. When executing INT2 quantization with only four representable levels, outliers dominate the scale factor, wasting much of its range on rare spikes and compressing normal values into just one or two effective levels, thus degrading attention quality substantially.

Rotation-based quantization addresses this by applying a fixed orthogonal transform, typically a Hadamard transform, to redistribute outlier energy across all channels. This approach works reasonably well at INT4. At INT2, a deeper problem remains; the rotation is data-oblivious. It can smooth activation ranges but does not know which directions the attention mechanism actually reads. Spreading quantization error uniformly is not the same as directing it into low-importance directions.

What OSCAR Does Differently

OSCAR’s key observation is that the rotation applied before quantization should be derived from attention statistics themselves — not from the raw distribution of KV activations.

For keys, the downstream error that matters is not the Euclidean reconstruction error of K but the error in attention logits. It is shown that this error is: ‖QK⊤ − QK̂⊤‖²F = tr((K − K̂)Q⊤Q(K − K̂)⊤). The weighting matrix is the query covariance Q⊤Q, not K⊤K. Directions where queries have large energy amplify quantization errors in logits. OSCAR estimates the empirical query covariance CQ = (1/N) Σ qn⊤qn from a calibration set, eigen-decomposes it, and uses the eigenvectors UQ as the key rotation basis.

For values, the relevant error is in the attention output SV, defined by how the attention score matrix S weights each value row. OSCAR uses the eigenvectors US of CS as the value rotation basis.

The final composed rotations are:

RK = UQ · HHad · Pbr RV = US · HHad · Pbr

Each of the three factors addresses a distinct failure mode of per-group low-bit quantization:

  • UQ / US aligns channels with attention-importance directions, diagonalizing the error-weighting matrix so the most important directions are identifiable.
  • HHad (Walsh-Hadamard transform) equalizes channel importance exactly. Lemma 1 proves every diagonal entry of HHad⊤ Λ HHad equals tr(Λ)/d — compressing the peaky eigenspectrum exposed by UQ to a uniform value across all channels.
  • Pbr (permuted bit-reversal) reorders channels to ensure each group receives one representative from each level of the importance hierarchy for any power-of-two quantization group size.

The research team provides Theorem 1 proving UQ and US are optimal under a frozen-error surrogate objective with diagonal residual assumptions.

The Serving System: Mixed-Precision Cache Layout

OSCAR integrates into SGLang’s production serving stack as an INT2 KV-cache mode with full compatibility with paged attention.

The KV cache layout uses three regions per request:

  • Sink tokens (first S0 = 64 tokens): stored in BF16 as attention sinks.
  • Recent tokens (last W = 256 tokens): stored in BF16 just before the current position.
  • History tokens (everything in between): stored as INT2 after OSCAR rotation and clipping.

At a 128K context length, the BF16 sink and recent windows represent only 0.24% of total tokens. The ablation shows (S=64, R=256) is the accuracy-efficiency knee, where smaller windows noticeably hurt accuracy and larger ones provide negligible additional benefit.

Write and read paths utilize fused Triton kernels. On the write path, each token is rotated, clipped to a calibration-derived percentile threshold, and quantized before being passed to the attention kernel. The value rotation RV is absorbed into the model’s projection weights offline, eliminating the runtime computational cost.

Outcome

The research team evaluated OSCAR on four model configurations: Qwen3-4B-Thinking-2507, Qwen3-8B, Qwen3-32B, and GLM-4.7-FP8 (358B parameters). Benchmarks include AIME25, GPQA-Diamond, HumanEval, LiveCodeBench v6, and MATH500, all at a maximum generation length of 32K.

Accuracy (at 2.28 bits per KV element):

Model BF16 Mean OSCAR Mean Gap to BF16
Qwen3-4B-Thinking-2507 75.64 71.86 −3.78
Qwen3-8B 70.84 69.42 −1.42
Qwen3-32B 74.19 74.17 −0.02
GLM-4.7-FP8 (358B) 77.89 78.16 +0.27

For context on competing methods: naive INT2 (no rotation) scores 0.00 on Qwen3-4B and Qwen3-8B. QuaRot-INT2 (Hadamard-only rotation) scores 1.40 on Qwen3-4B and 10.14 on Qwen3-8B. TurboQuant at 3.25 bits drops 43.90 points on Qwen3-4B-Thinking. Saw-INT4 at 4.25 bits reaches 73.11 on Qwen3-4B — OSCAR at 2.28 bits achieves 71.86.

The research team also compared against channel-wise methods on AIME25, showing results superior to KIVI-KV2*.

Long-context robustness (RULER-NIAH):

Model Method 16K 32K 64K 128K
Qwen3-4B-Thinking BF16 99.7 99.3 85.3 81.0
Qwen3-4B-Thinking OSCAR 97.8 87.6 61.9 39.5
Qwen3-8B BF16 98.9 97.3 79.2 78.2
Qwen3-8B OSCAR 93.9 86.3 61.9 45.0

For throughput (H100) at a 100K context, job-level throughput reaches 6.17× over BF16 on Qwen3-4B-Thinking and 7.83× on GLM-4.7-FP8.

Marktechpost’s Visual Explainer

OSCAR (Offline Spectral Covariance-Aware Rotation) is a 2-bit KV cache quantization system designed for long-context LLM serving.

Setup Requirements:

  • Hardware: NVIDIA H100 GPU recommended.
  • Install SGLang:
    pip install sglang[all] --upgrade
    pip install triton
    

Download Pre-Computed Rotations via RotationZoo:

from modelscope import snapshot_download

# Download RotationZoo for your model
rotation_path = snapshot_download(
    'togethercomputer/OSCAR-RotationZoo'
)

Inference Example:

from openai import OpenAI

client = OpenAI(
    base_url='http://localhost:30000/v1',
    api_key='none'
)

response = client.chat.completions.create(
    model='Qwen/Qwen3-8B',
    messages=[{"role": "user","content": "Your long-context prompt here"}],
    max_tokens=1024
)
print(response.choices[0].message.content)

Key Takeaways

  • OSCAR quantizes LLM KV caches to 2-bit precision, achieving near-BF16 accuracy and full compatibility with paged KV-cache serving.
  • Significant reductions in KV memory and improvements in decode speed and throughput are demonstrated.
  • Pre-computed rotation matrices are available for select models , eliminating the need for recalibration.