What drives home value: a scorecard for consistent underwriting
I wanted to know which factors really drive home value. A Random Forest explained 82% of it, and three different methods agreed on the top driver.
R² on held-out data, above the 0.75 target
average error (RMSE), down from $71,734
attribution methods ranked median income first
modeled yearly savings for a 50-deal fund
Why I built it
I work in real estate investment, so valuation isn’t abstract to me. Two analysts can look at the same property and land 10 to 20 percent apart, because each one weighs location, age and income a little differently. Nobody’s wrong, exactly. There’s just no shared, documented answer to which factors matter most. That’s the question I wanted this model to answer.
What I worked with
20,640 California neighborhoods from the 1990 census. It’s old data, and each row is a neighborhood, not a house, so I treated this as a test of the method, not a pricing tool. Two things jumped out right away: prices climb with income, and they hug the coast.
Cleanup came first. Four of the raw columns (total rooms, bedrooms, population and households) were basically measuring the same thing, with correlations above 0.85, so I turned them into three ratios that actually mean something: rooms per household, bedrooms per room and people per household. 207 rows were missing bedroom counts, about 1%, and I filled those with the median for their distance from the ocean. 965 homes (4.7%) were capped at $500,001, which I flagged instead of pretending it wasn’t there. Then an 80/20 split, stratified by income.
Setting the bar first
Before training anything, I set the bar: an R² of at least 0.75 on data the model had never seen. Then I ran two models. Linear regression, because anyone can read its coefficients, and a 200-tree Random Forest, because coastal pricing doesn’t follow a straight line.
The Random Forest cleared the bar; linear regression didn’t
| Model | Test RMSE | Test MAE | Test R² |
|---|---|---|---|
| Linear regression | $71,734 | $51,397 | 0.611 |
| Random Forest | $48,720 | $31,602 | 0.821 |
The linear model missed. The Random Forest cleared it and cut the average error by about $23,000. That gap is the coastline.
What actually drives value
I didn’t want to trust a single method, so I checked three: the linear coefficients, the Random Forest’s built-in importance and permutation importance. All three put median income first. In the linear model, one standard deviation more income is worth about $77,458 in predicted value.
| Feature | Linear | Random Forest | Permutation | Average rank |
|---|---|---|---|---|
| Median income | 1 | 1 | 1 | 1.00 |
| Longitude | 2 | 4 | 4 | 3.33 |
| Latitude | 3 | 5 | 2 | 3.33 |
| Inland location | 4 | 2 | 5 | 3.67 |
| People per household | 8 | 3 | 3 | 4.67 |
| Housing median age | 6 | 6 | 6 | 6.00 |
| Rooms per household | 7 | 7 | 7 | 7.00 |
The methods didn’t agree on everything. People per household ranks eighth in one and third in the other two. I kept both answers, because when methods split like that, it usually means a feature matters through how it interacts with the others rather than on its own.
How I’d use it
For an acquisitions team, this becomes a scorecard: screen every inbound deal on the few drivers that matter most, then spend analyst time on the ones that pass. For underwriting, it’s a documented reason behind every weight.
The notebook puts numbers on that for a California fund doing 50 acquisitions a year at $400,000 each. Rule-of-thumb underwriting runs about a 15% pricing error, or $60,000 a deal. The Random Forest’s average error is $31,602, about 8%. That’s $28,398 a deal, roughly $1.42 million a year in avoided overpayment, plus around 400 analyst hours saved by screening on the top three features. To be clear, that’s a modeled business case built on assumptions, not a production result.
What I’d fix before trusting it
- Current data. 1990 data proves the method. A real version needs today’s data and a refresh schedule.
- The $500,001 cap. It hides the top of the market and skews the rankings there.
- Neighborhoods, not houses. The weights carry over to property-level work. The predicted prices don’t.
- Fair lending. Location is fine for segmenting a market. It’s risky the moment it touches a price a borrower sees, so that boundary needs to be written down and reviewed.
- Time. This tells you what a place is worth, not what it will be worth. Projecting returns needs a separate appreciation model.