Quantization risk tool — feedback

This is optional and there is no wrong amount to fill in. A single number is useful; so is a sentence saying the tool was not much use.

Optional. Anything that fits - "student ML researcher at Berkeley", "ML engineer", "just tinkering". Applies to either section.


1. A measured result, if you have one

Only if you happened to run an eval. Skip the whole section otherwise.

w4a16 / w8a8 / w8a16 / fp8 / fp8-dynamic / nvfp4 / other

e.g. 1.5 - a number in billions

e.g. Qwen2.5, Llama-3.1. Leave blank if you cannot share it

mmlu / arc_challenge / hellaswag / gsm8k / winogrande / truthfulqa / other

unquantized score, 0-100 scale

quantized score, 0-100 scale

e.g. lm-eval-harness, MMLU 5-shot. Needed to compare like with like

e.g. llm-compressor GPTQ, 512x2048 calibration, group 128

especially: did anything go wrong, or was this a config you would not recommend?

only if you want a reply


2. Open feedback

Send this on its own if you like — nothing above is required.

What was confusing, what you expected it to do, whether it was worth your time. Criticism is more useful than praise here.

Submissions are reviewed by a human before anything enters the dataset; nothing is ingested automatically. Only what you type here is sent — there is no tracking on this page or in the tool.