Inference Benchmarks
Compare KV cache performance and speculative decoding on GPU
KV Cache
Speculative Decoding
Configuration
Prompt
Max New Tokens
Run KV Cache Benchmark