arrow_backRetour aux issues
poojithdevan4D/pooji-vllm
#6
Débutant
Ouvrirarrow_forward
Débutant
Ouvrirarrow_forward
Débutant
Ouvrirarrow_forward
Working INT8/INT4 weight-only quantization
ecoDébutant
help wanted
performance
descriptionDescription
Decode is bandwidth-bound: `benchmarks/roofline.py` measures 160 GB/s achievable against 942 MiB of weights, a 6.22ms/step floor. Halving weight bytes should roughly halve decode time. This is the largest single win available.
**Prior attempt (failed, documented in BENCHMARKS.md §10):** torchao `Int8WeightOnlyConfig` made it 2.5x *slower* (11.5 → 29.1 ms/step) and produced garbage output. Eager mode dequantizes to fp16 before the matmul, which adds traffic rather than removing it, and the tensor subclasses do not survive our hand-written forward pass.
**What is actually needed:** a fused dequant-matmul (Marlin-style) so weights are read as int8/int4 and expanded inside the kernel.
**Done when:** measured ms/step improves at B=1..16 and output quality is checked, not assumed.
Issues similaires
calkit/calkit
star53
Poids du dépôt moyen
VS Code extension should be robust to YAML parser errors
Seeing this error: ``` Failed to read calkit.yaml: YAMLParseError: A block sequence may not be used as an implicit map…
Python
bug
good first issue
fu351/Doberman-Core
star211
Poids du dépôt léger
dash: a manual Refresh control
The dashboard polls: `refreshStats()` (`src/doberman/dash/app.py:408`) every 5 s and `refreshPending()` (`:546`) every …
Python
enhancement
good first issue
fu351/Doberman-Core
star211
Poids du dépôt léger
dash: "Copy details" button on each pending-approval card
Each pending-approval card in the dashboard (`renderPending`, `src/doberman/dash/app.py:448-544`) shows the risk badge,…
Python
enhancement
good first issue