Research Papers

The papers behind the cost curve — attention variants, KV-cache compression, quantization and serving efficiency.

What this covers

Key papers on inference efficiency, attention architectures, quantization, caching and the economics of large-model serving.

Why it matters

Tokenomy tracks AI economics end to end — pricing, throughput, quality and energy. This library page keeps the underlying evidence one click from the tools and research that use it.