arrow_backRetour aux issues
Samarjamal326/MediOrchesctrator-Agent
#20
Débutant
Ouvrirarrow_forward
Débutant
Ouvrirarrow_forward
Débutant
Ouvrirarrow_forward
Medical Data Preprocessing and Cleaning Pipeline
ecoDébutant
good first issue
data-science
descriptionDescription
## Objective
Create a preprocessing and cleaning pipeline to prepare raw medical datasets for knowledge retrieval.
## Tasks
- [ ] Develop scripts to parse and clean raw medical documents (PDFs, text, CSVs)
- [ ] Implement text normalization, noise removal, and deduplication
- [ ] Validate document structure and output clean, structured text files
- [ ] Document data cleaning guidelines and output validation steps
## Dependencies
- Depends on #8 (Data Collection)
Issues similaires
calkit/calkit
star53
Poids du dépôt moyen
VS Code extension should be robust to YAML parser errors
Seeing this error: ``` Failed to read calkit.yaml: YAMLParseError: A block sequence may not be used as an implicit map…
Python
bug
good first issue
fu351/Doberman-Core
star211
Poids du dépôt léger
dash: a manual Refresh control
The dashboard polls: `refreshStats()` (`src/doberman/dash/app.py:408`) every 5 s and `refreshPending()` (`:546`) every …
Python
enhancement
good first issue
fu351/Doberman-Core
star211
Poids du dépôt léger
dash: "Copy details" button on each pending-approval card
Each pending-approval card in the dashboard (`renderPending`, `src/doberman/dash/app.py:448-544`) shows the risk badge,…
Python
enhancement
good first issue