I Tested ChatGPT on ARV. Here's Where It Failed
TLDRChatGPT gave me plausible ARV answers and found some useful comps, but it changed methods between runs and missed a stale listing and a flood-zone problem. I use it to build a first-pass comp list, then I verify the evidence and make the call myself.
Table of Contents
- What I Tested
- Round One: A Good-Looking Answer With a Different Method
- Round Two: The Misses That Mattered
- Training a Custom GPT
- A Source-Faithful AI Workflow
- FAQ
What I Tested
Can ChatGPT calculate the after repair value of a house well enough to help an investor?
My answer after testing it is yes, but only as a research assistant. It can surface candidate comps, show the math, and get you moving faster. It can also return a polished answer while missing the one fact that ruins the deal.
I had 14 years of buying real estate behind me when I ran this test. I used actual properties I knew, including one of my own listings and another house near the edge of a flood zone. That gave me a way to check the answer against facts I already understood.
I ran three versions of the test:
| Test | What I Asked | What Happened |
|---|---|---|
| Out of the box | Find the ARV of a property | It found useful comps and landed at $250,000 |
| Fresh chat | Run the same property again | It returned $260,000, but used a different valuation method |
| Custom GPT | Follow my comping rules | It improved the work, then still missed a stale listing and the flood-zone issue |
The $10,000 difference between the first two answers did not bother me much. ARV is a range. The method changing underneath the number bothered me a lot.
The number looked believable. I still needed to see how the machine got there.
Round One: A Good-Looking Answer With a Different Method
The first response did four things I liked. It selected three to five properties in what looked like the right neighborhood. It converted each sale to price per square foot. It gave me a red-flag list. Then it settled on a $250,000 ARV.
That is a useful start.
I opened a new chat so it could not lean on the first answer and asked the same question again. This time it returned $260,000. Again, that difference was not the real problem. On the second run it leaned on automated estimates from property websites, averaged those estimates, and added what it guessed a renovation might contribute.
That is not the same as building a sales-comparison case from verified sold properties. It was a different method wearing the same confident tone.
For a flip, I am trying to estimate what the renovated house can sell for. Today’s automated estimate may describe the house in its current condition, and sometimes not even that very well. It does not tell me what a specific renovation will produce.
The offer calculation is a separate step. In the video, my quick example used the 70 percent rule: a $300,000 ARV times 70% is $210,000, then subtracting a $60,000 rehab gives a $150,000 target purchase price. That is one of my fast screening rules, not a substitute for checking the real costs and exit.
Common MistakeDo not accept an ARV because the number looks reasonable. Ask for the sold properties, open the source records, and check whether the method stayed the same.
Round Two: The Misses That Mattered
For the next round I used a nearby house for a reason. I knew the neighborhood, I knew what was sitting on the market, and I knew the property touched the edge of a 100-year flood zone.
ChatGPT returned a $325,000 ARV. I saw three problems.
First, no house in that neighborhood had sold above $300,000 in the recent period I was looking at. The subject was larger than many nearby houses, but the AI was asking it to break through a ceiling the market had not proved.
Second, it missed my renovated listing nearby. That house was offered at $275,000 and was not selling. I had received one low offer around $200,000. An active listing is not a sold comp, but a stale listing can disprove an aggressive ARV. If buyers will not pay $275,000 for the renovated house already available, I need a very good reason to believe the next one will sell for $325,000.
Third, it missed the flood-zone issue. In a separate question, ChatGPT could explain why flood insurance and buyer resistance may reduce value. It had the general information. It did not connect that information to the property in the ARV test.
That was the weird part. It knew the flood-zone fact when I asked directly. It just did not bring that fact into the deal in front of it.
The Expensive MissA model can get most of the comp work right and still miss the one condition that changes the exit. Flood maps, zoning, easements, setbacks, and other deal-specific facts need a separate check.
The arithmetic was not the problem. The missing evidence was.
Training a Custom GPT
I did not quit after two rounds. I fed ChatGPT the transcript from my full manual comping lesson and asked it to turn those rules into instructions for a custom GPT.
The instructions told it to:
- Find three to five sold comps that match on features, sale date, and neighborhood.
- Check those properties for my DOA red flags.
- Show the price-per-square-foot math for each comp.
- Review current and pending listings that might support or contradict the comp set.
The first version did not work well. Neither did the second. I kept adjusting it, probably three, four, or five times, before the answers became useful.
Even then, it missed the flood zone again. It also missed the stale listing that contradicted its value.
So training helped, but it did not turn the tool into an independent underwriter. It made the assistant better at following my checklist. I still had to know the checklist, inspect the sources, and catch the gaps. I am not smart enough to remember every odd factor on command; the checklist is what keeps me from pretending I did.
If I do not know how to comp the house, I cannot tell when the AI is wrong.
A Source-Faithful AI Workflow
The source supports using AI for the slow first pass, not the final decision. The sequence below turns the checks shown in the test into a practical workflow; it is not a claim that Ross recited these eight steps as a named system.
Here is the sequence:
- Ask for three to five sold comps and require a link or source record for each one.
- Ask it to show the sale price, square footage, price per square foot, sale date, and why each property is comparable.
- Run the same property in a fresh chat and compare the method, not just the final number.
- Open every comp and verify the basic facts myself.
- Check the neighborhood boundaries with a map, a census-tract view, and what I know from being there.
- Run a separate property check for flood zones, parcel issues, zoning, setbacks, low ceilings, and odd construction.
- Look at active and pending listings for evidence that challenges the sold-comp range.
- Set my own ARV range and use that in my offer math.
The full manual method lives in How to Calculate ARV Like a Pro. That is the knowledge you need before AI becomes useful.
This is also why I use IMBY. Neighborhood boundaries are not just marks on a screen. I want to know the place well enough to feel when a comp crossed a major road, railroad, or other boundary into a different market.
AI saves me time because it can do several minutes of searching and organize the result. I can reject a bad comp and ask for another one. I can make it show its work. I just do not hand it the thinking.
AI can do the legwork. I still have to own the number.
FAQ
Can ChatGPT calculate ARV by itself?
It can produce an estimate, but I would not buy a house from that answer alone.
Why did the $250,000 and $260,000 answers worry you?
The $10,000 spread did not worry me. The switch in method did. One run built something close to a sold-comp analysis, while the next leaned on automated estimates and guessed at the effect of improvements. If the method moves around, the answer is harder to trust.
Does a custom GPT solve the problem?
It makes the output more consistent because you can give it your rules. Mine still missed a flood-zone issue and a stale listing after several rounds of training. A custom GPT is a better assistant, not permission to stop verifying.
What should I make ChatGPT show me?
Make it show every comp, the source link, sale date, square footage, price per square foot, neighborhood reason, and any active listing that challenges the answer. Then open the records. If you cannot check the evidence, you do not have an ARV you can defend.
I am just starting out. Should I use AI for comps?
Yes, as long as you do the manual work beside it. Build your own comp range, compare it with the AI output, and investigate every difference. Then run another property and do it again. That is how the tool saves time without stealing the reps you need.