Where the others stand
Every other layer stops at tool output. None of them curates thinking.
Only one of them (RTK) has been graded on whether the agent still solves the task, and it lost. Their headline numbers are single-session or single-scenario token counts; none reports task outcomes. Every Parsec number below is measured on one fixed model with the official SWE-bench harness, or on real agent traces with the same needed label the curator trains on.
Capability matrix · Parsec vs the field
Every other layer rewrites, summarizes, or deletes what your agent reads.
| Parsec | Caveman | Headroom | Woz | RTK | |
|---|---|---|---|---|---|
| Compresses tool output | yes | yes | yes | yes | shell only |
| Compresses thinking | yes | ✗ no | ✗ no | ✗ no | ✗ no |
| Trims the tool roster every turn | yes | ✗ no | ✗ no | ✗ no | ✗ no |
| Never alters readable messages | yes | ✗ no | ✗ no | ✗ no | ✗ no |
| Kept content byte-exact (no rewording) | yes | ✗ no | ✗ no | ✗ no | ✗ no |
| Learned from real agent traces | yes | ✗ no | ✗ no | ✗ no | ✗ no |
| Measured on the wire, not estimated | yes | provider-rep. | estimated | estimated | isolated |
| Graded on task solves (SWE-bench, /100) | 62 | 58 | 58 | 55 | 54 |
| Cuts output tokens | −43% | −30% * | −2% | −1% | +8% |
| Cuts input tokens (cache-aware) | −54% | −19% | +6% | −35% | +16% |
| Cost per solved task | $1.45 | $2.05 | $3.66 | $2.33 | $3.07 |
| Mean added latency per call | 6.2 s | 6.9 s | 8.0 s | 14.3 s | 9.3 s |
| Runs local | yes | yes | yes | ✗ no | yes |
* Caveman cuts output by rewriting the reply into caveman-speak, not by reducing reasoning. Token, cost, time and solve columns are the six-arm SWE-bench Verified run (same model, official grader). Competitor claims as published on their sites and READMEs, 2026-09-09.
Rows = input types · columns = layers
Everyone shrinks some inputs. Only Parsec drops the thinking and the roster — and never a word you can read.
| Parsec | Headroom | RTK | Caveman | Woz | |
|---|---|---|---|---|---|
| File reads | 82.2% | 47% | — not reported | — not reported | — not reported |
| Command output | 68.1% | — not reported | 89% | — not reported | — not reported |
| Search results | 98.2% | 92% | — not reported | — not reported | — not reported |
| Web content | 67.9% | — not reported | — not reported | — not reported | — not reported |
| Context & plans | 84.9% | — not reported | — not reported | — not reported | — not reported |
| Thinking | 69.7% | 0 | 0 | 0 | 0 |
| Tool roster | 77% | 0 | 0 | 0 | 0 |
| Your messages | 0 | 0 | 0 | — not reported | 0 |
| Their own headline | 78.4% | 87% | 89% | 33% | 35% |
| Same meter — input tokens cut, SWE-bench | −54% | +6% | +16% | −19% | −35% |
Two rows have exactly one nonzero cell — thinking and the tool roster — both Parsec. On the same-meter row Parsec is −54%; Headroom and RTK add tokens.
Each competitor cell is that vendor's own published number, computed its own way (RTK's 89% is command output only — Read, Grep and Glob bypass its hook). Parsec cells are the share of that input type the agent never reused. The last row is one meter: provider-reported input tokens, same tasks, same model.
SWE-bench Verified · cost per task · Parsec vs no compression
Cost per task, every case.
Parsec drops the tool output and thinking the agent never uses. Across 100 SWE-bench Verified tasks that cut cost 39%; 21 cost more.
| SWE-bench Verified case | No compression | Parsec | Change |
|---|---|---|---|
| matplotlib-20488 | $3.65 | $0.52 | −86% |
| matplotlib-20826 | $4.11 | $1.22 | −70% |
| django-11555 | $1.95 | $0.76 | −61% |
| sympy-22456 | $6.24 | $2.74 | −56% |
| django-15037 | $1.36 | $1.92 | +41% |
| sympy-23413 | $3.06 | $4.79 | +57% |
| psf/requests-5414 | $0.31 | $0.61 | +96% |
| All 100 | $147.30 | $89.65 | -39% |
Parsec solved 12 the baseline missed; the baseline solved 7 Parsec missed. Net +5 solves at 39% lower total cost. We show the red rows.
The curator drops what the agent never reuses
Whatever the workload, ~three-quarters is never reused
Parsec's curator model scores every chunk of tool output and thinking and drops what the agent never reuses. The share it can drop lands near three-quarters no matter the workload.
| Workload | Traces | Context in | Sent | Saved |
|---|---|---|---|---|
| Software engineering | 3,363 | 131,581,592 | 25,765,837 | 80.4% |
| Writing & docs | 1,767 | 82,695,132 | 17,873,238 | 78.4% |
| Data analysis | 1,436 | 60,466,980 | 17,037,932 | 71.8% |
| Research / web | 509 | 10,731,243 | 2,649,113 | 75.3% |
| Devops | 706 | 37,706,700 | 11,077,929 | 70.6% |
| Other | 20,706 | 615,429,868 | 128,598,905 | 79.1% |
| All workloads | 28,487 | 938,611,515 | 203,002,954 | 78.4% |
Measured on 28,487 real agent traces.