Research Papers
The papers behind the cost curve — attention variants, KV-cache compression, quantization and serving efficiency.
What this covers
Key papers on inference efficiency, attention architectures, quantization, caching and the economics of large-model serving.
Why it matters
Tokenomy tracks AI economics end to end — pricing, throughput, quality and energy. This library page keeps the underlying evidence one click from the tools and research that use it.