Building machine learning models with Claude Code
Claude Code is my everyday tool. Here’s how I split the work: what it does, and what I own at every step.
Why I use it
Claude Code is my everyday tool. It works inside my own files and command line: it reads data and code, writes and runs scripts, checks the results and keeps going across multi-step tasks. The skill isn’t typing code. It’s directing it toward a business outcome, and that’s what lets me build the thing I’m proposing instead of just describing it.
With it, I can go from a business question to a tested, documented model in one working session. A second agent reviews every line of code, every project is documented in Linear, Obsidian and GitHub, and I own the judgment, the validation and the sign-off.
Who does what
| Stage | What Claude Code does | What I own |
|---|---|---|
| 1. Frame the decision | Drafts the problem statement, success criterion and cost of errors from my brief | The decision, the metric and the threshold |
| 2. Prepare the data | Profiles the data, flags missing values, outliers and leakage, and writes the cleaning code | Imputation and exclusion rules, data access and privacy |
| 3. Build and compare | Trains a baseline and challengers, tunes with cross-validation and compares them on AUC or RMSE | Model choice against the cost of errors, and a bias review of the features |
| 4. Explain | Produces feature attribution, ROC and diagnostic charts, and plain-language findings | Confirming the interpretation with the business owner |
| 5. Deploy | Wraps the model in a Streamlit app or a scheduled script and writes tests | Release approval, monitoring and retraining triggers |
| 6. Review | A separate AI agent audits and reviews all of the code before anything is used | Acting on what the review finds, and final approval |
| 7. Document | Writes up assumptions, limits and handover notes in Linear, Obsidian and GitHub | Sign-off for governance and independent validation |
Every case study here follows that order: the business question first, the bar set before training, several models against a baseline, and the metric matched to the cost of being wrong.
A second agent checks the work
I don’t let the agent that wrote the code be the only one that checks it. A separate agentic reviewer audits and reviews all of the code: it reads what was built, tests it against what was asked for, and flags anything that looks wrong before it goes anywhere near a team.
Every project is documented
A project nobody else can understand isn’t finished. I document every project where other people can find it, pick it up and build on it, using three tools that each do one job:
- Linear. The project itself: scope, milestones, owners, status and action items, connected to our engineering team.
- Obsidian. The knowledge behind it: notes, decisions, assumptions and the reasoning for why it was built the way it was.
- GitHub. The code and its full history, so every change is tracked and anyone can review it or roll it back.
That’s what makes the work collaborative instead of living on one laptop, and it’s how a portfolio of 70+ projects stays organized.
Where it fits in an AI program
The same workflow builds scoring models for an AI project intake queue, risk and propensity models, anomaly detection on operational data, search across policy documents, and internal dashboards. Each one replaces a manual review or a slow handoff, and I can prototype it in days. That’s how I keep working tools in front of portfolio owners early and save engineering time for production. It’s also why I keep a Claude skills library: once a workflow works, it should be reusable.
The guardrails
- Validation. Every model gets checked on held-out data against a bar set before training.
- Review. An agentic reviewer audits and reviews all code, and it’s tested before anyone uses it.
- Version control. Code lives in GitHub, the project lives in Linear and the documentation lives in Obsidian, so anyone on the team can pick it up.
- Data. Sensitive data stays in approved environments.
- Labels. Prototypes stay labeled as prototypes until they pass the same validation and change control as any model that influences a decision.
- Documentation. Assumptions, limits and bias exposure get written down and reviewed with risk and compliance partners.