arrow_backRetour aux issues
poojithdevan4D/pooji-vllm
#5
Débutant
Ouvrirarrow_forward
Débutant
Ouvrirarrow_forward
Débutant
Ouvrirarrow_forward
Chunked prefill (stop TTFT growing with queue position)
ecoDébutant
help wanted
performance
descriptionDescription
Prefill currently runs one request at a time in its own tick. With 8 concurrent clients, measured TTFT climbed 47ms → 328ms purely by queue position.
**Where:** `pooji_vllm/llm_engine.py` — `step()`, the `while self.waiting:` block.
**Approach:** split a long prompt into chunks of N tokens and mix those chunks into the same batch as decode tokens, instead of running a dedicated prefill pass. This requires `_forward` to accept a batch with different T per request (ragged), which is the main design work.
**Done when:** TTFT under 8 concurrent clients is roughly flat, with no throughput regression on `benchmarks/bench.py`.
Issues similaires
calkit/calkit
star53
Poids du dépôt moyen
VS Code extension should be robust to YAML parser errors
Seeing this error: ``` Failed to read calkit.yaml: YAMLParseError: A block sequence may not be used as an implicit map…
Python
bug
good first issue
fu351/Doberman-Core
star211
Poids du dépôt léger
dash: a manual Refresh control
The dashboard polls: `refreshStats()` (`src/doberman/dash/app.py:408`) every 5 s and `refreshPending()` (`:546`) every …
Python
enhancement
good first issue
fu351/Doberman-Core
star211
Poids du dépôt léger
dash: "Copy details" button on each pending-approval card
Each pending-approval card in the dashboard (`renderPending`, `src/doberman/dash/app.py:448-544`) shows the risk badge,…
Python
enhancement
good first issue