AI Knowledge · Case study 07 · Agentic development

Building machine learning models with Claude Code

Claude Code is my everyday tool. Here’s how I split the work: what it does, and what I own at every step.

Why I use it

Claude Code is my everyday tool. It works inside my own files and command line: it reads data and code, writes and runs scripts, checks the results and keeps going across multi-step tasks. The skill isn’t typing code. It’s directing it toward a business outcome, and that’s what lets me build the thing I’m proposing instead of just describing it.

With it, I can go from a business question to a tested, documented model in one working session. A second agent reviews every line of code, every project is documented in Linear, Obsidian and GitHub, and I own the judgment, the validation and the sign-off.

Who does what

StageWhat Claude Code doesWhat I own
1. Frame the decisionDrafts the problem statement, success criterion and cost of errors from my briefThe decision, the metric and the threshold
2. Prepare the dataProfiles the data, flags missing values, outliers and leakage, and writes the cleaning codeImputation and exclusion rules, data access and privacy
3. Build and compareTrains a baseline and challengers, tunes with cross-validation and compares them on AUC or RMSEModel choice against the cost of errors, and a bias review of the features
4. ExplainProduces feature attribution, ROC and diagnostic charts, and plain-language findingsConfirming the interpretation with the business owner
5. DeployWraps the model in a Streamlit app or a scheduled script and writes testsRelease approval, monitoring and retraining triggers
6. ReviewA separate AI agent audits and reviews all of the code before anything is usedActing on what the review finds, and final approval
7. DocumentWrites up assumptions, limits and handover notes in Linear, Obsidian and GitHubSign-off for governance and independent validation

Every case study here follows that order: the business question first, the bar set before training, several models against a baseline, and the metric matched to the cost of being wrong.

A second agent checks the work

I don’t let the agent that wrote the code be the only one that checks it. A separate agentic reviewer audits and reviews all of the code: it reads what was built, tests it against what was asked for, and flags anything that looks wrong before it goes anywhere near a team.

Every project is documented

A project nobody else can understand isn’t finished. I document every project where other people can find it, pick it up and build on it, using three tools that each do one job:

  • Linear. The project itself: scope, milestones, owners, status and action items, connected to our engineering team.
  • Obsidian. The knowledge behind it: notes, decisions, assumptions and the reasoning for why it was built the way it was.
  • GitHub. The code and its full history, so every change is tracked and anyone can review it or roll it back.

That’s what makes the work collaborative instead of living on one laptop, and it’s how a portfolio of 70+ projects stays organized.

Where it fits in an AI program

The same workflow builds scoring models for an AI project intake queue, risk and propensity models, anomaly detection on operational data, search across policy documents, and internal dashboards. Each one replaces a manual review or a slow handoff, and I can prototype it in days. That’s how I keep working tools in front of portfolio owners early and save engineering time for production. It’s also why I keep a Claude skills library: once a workflow works, it should be reusable.

The guardrails

  • Validation. Every model gets checked on held-out data against a bar set before training.
  • Review. An agentic reviewer audits and reviews all code, and it’s tested before anyone uses it.
  • Version control. Code lives in GitHub, the project lives in Linear and the documentation lives in Obsidian, so anyone on the team can pick it up.
  • Data. Sensitive data stays in approved environments.
  • Labels. Prototypes stay labeled as prototypes until they pass the same validation and change control as any model that influences a decision.
  • Documentation. Assumptions, limits and bias exposure get written down and reviewed with risk and compliance partners.
More case studies

Keep reading

Back to AI Knowledge