1

Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding
Reasoning language models have demonstrated remarkable capabilities on challenging tasks by generating elaborate chain-of-thought (CoT) …
TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval
Retrieval-augmented generation (RAG) extends large language models (LLMs) with external data sources to enhance factual correctness and …
Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
  • Use activation and weight low-bit quantization to improve throughput.
Atom: Low-bit Quantization for Efficient and Accurate LLM Serving