Frontier: A Discrete-Event Simulator for Modern LLM Serving
-
Updated
Sep 27, 2026 - Python
Frontier: A Discrete-Event Simulator for Modern LLM Serving
An open toolkit and public dataset hub for collecting, sanitizing, analyzing, and visualizing coding agent traces.
A high-performance, universal serving framework for any-to-any models.
multi-robot serving engine for cloud robot foundation models
Learn the ins and outs of efficiently serving Large Language Models (LLMs). Dive into optimization techniques, including KV caching and Low Rank Adapters (LoRA), and gain hands-on experience with Predibase’s LoRAX framework inference server.
CLI toolkit for LLM inference preflight, vLLM serving configuration, benchmarking, telemetry, and capacity analysis.
Discrete-event simulator for studying throughput, TTFT, and KV-cache transfer tradeoffs in disaggregated LLM serving.
Simulating KV cache reuse across multi-turn LLM conversations using real ShareGPT traces. Compares sticky, least-load, hybrid, and cost-aware routing strategies for cache locality vs load balance.
Discrete-event simulation of SLO-aware autoscaling for multi-instance LLM serving — comparing reactive, predictive, and conservative policies across workload patterns and startup delays.
Discrete-event simulation of LLM request routing across multiple serving instances: round-robin, least-load, prefix-aware, and hybrid cache-aware routing with queue-depth threshold sweep.
Benchmarking output length prediction quality and calibration for KV-cache-aware LLM serving scheduling. A noisy predictor with 1.5x safety margin captures 99.7% of oracle gains — calibration matters more than accuracy.
Comparing KV-cache-aware scheduling policies for LLM serving: greedy preempts 49% of requests under memory pressure while memory-first achieves 43% higher effective throughput with zero preemptions.
Comparing LLM serving under real ShareGPT traces vs synthetic Poisson workloads. Real distributions produce 42x worse P99 TTFT and earlier SLO violations — scheduling policies only differentiate under heavy-tailed real traffic.
Step-level simulation of chunked prefill under KV cache pressure. Block policy wastes 89.6% of prefill chunks while achieving the same throughput as upfront rejection. Chunk waste rate is the metric standard throughput metrics miss.
To associate your repository with the serving-infrastructure topic, visit your repo's landing page and select "manage topics."