April 29, 2026 · papers

Letting a language model tune the energy knobs

The runtime parameters that decide what inference costs in energy are numerous and they interact, so in practice nobody tunes them by hand. Katelyn Crumpacker's framework puts a chat-based LLM in the loop with structured feedback prompts. The enhanced prompt template converges in 3.4 prompts on average against 5.2 for the baseline, beats Sobol sampling on convergence speed, and reaches lower final energy per token across hardware setups.

← All news