Latency, diversity, and quality benchmark for autoregressive decoding strategies on RTX 2070: greedy, top-k, top-p, min-p, and beam search across GPT-2 models.
benchmarking text-generation pytorch top-k beam-search sampling perplexity gpt2 greedy-decoding top-p llm-inference decoding-strategies min-p mlsystems
-
Updated
Jul 19, 2026 - Python