# Research Papers

> Key papers on inference efficiency, attention architectures, quantization, caching and the economics of large-model serving.

Source: https://tokenomy.ai/research/research-papers
Last updated: 2026-09-21
Publisher: Tokenomy — FinOps for AI
License: free to quote with attribution and a link to the source URL.

The papers behind the cost curve — attention variants, KV-cache compression, quantization and serving efficiency.

## What this covers

Key papers on inference efficiency, attention architectures, quantization, caching and the economics of large-model serving.

## Why it matters

Tokenomy tracks AI economics end to end — pricing, throughput, quality and energy. This library page keeps the underlying evidence one click from the tools and research that use it.

## Related pages

- [Research hub](https://tokenomy.ai/research)
- [AI News Hub](https://tokenomy.ai/research/ai-news-hub)
- [Model Benchmarks](https://tokenomy.ai/research/model-benchmarks)
- [Research Papers](https://tokenomy.ai/research/research-papers)
- [Innovation Tracker](https://tokenomy.ai/research/innovation-tracker)
- [Conference Calendar](https://tokenomy.ai/research/conference-calendar)
- [Tiktoken Guide](https://tokenomy.ai/research/tiktoken-guide)
- [Free tools](https://tokenomy.ai/tools)
