Accuracy, with its limits
How accurate is it?
The honest answer is that accuracy is not one number, and anybody who gives you one without saying what it was measured on is selling you something. Here is what Guru measures, how, and where it stops.
Scan a propertyAccuracy depends on the evidence, so we tell you which you got
An estimate for a three-bedroom house in a suburb with twelve recent comparable sales and a an estimate in a suburb with one are not the same product, and pretending otherwise is where most automated valuation tools quietly go wrong. Guru does not average that difference away. Every report names the route its number came from and carries a confidence score earned from that route.
Which means the useful question is not "how accurate is Guru" but "how accurate is an estimate built on evidence like mine". That is a question with an answer, and the rest of this page is about how we get it.
The outcome ledger
Every estimate Guru publishes is written to a ledger, together with a fingerprint of the evidence behind it: which method won, how many comparables it rested on, whether any were confirmed sales, which size basis was used, which metro. The estimate and the reasoning are recorded together, before anyone knows the answer.
Then reality arrives. The property sells, or it is relisted at a new price, or it is quietly withdrawn. Each of those is an observation, and each is scored against the prediction that preceded it: how far off, in which direction, and whether the real price landed inside the published range. Predictions written after a sale are not scored against it, because they were not predictions.
Those scores then set future ranges. Instead of a confidence band chosen by assumption, the question becomes empirical: for estimates that rested on evidence of this exact shape, where did real prices actually land? The measured spread becomes the band.
Where that stands today. The ledger records every estimate from day one, and the scoring runs nightly. But a category of evidence must reach 30 scored outcomes before it is allowed to set a range, so on most reports today the band still comes from the conservative default rather than from measurement, and the report records which one it used. That is the honest position: the machinery is built and running, the measurement arrives with volume, and we would rather tell you that than let an architecture diagram read as a result.
The rules that keep the scoring honest
- A bucket must earn the right to speak. Fewer than 30 scored outcomes and it is noise, and a "measured" band built on eleven sales is a worse lie than an honest default because it looks empirical. Below the threshold Guru falls back to a conservative assumption and says so.
- Portal "sold" prices are not sale prices. A sold card generally carries the last asking price, which is an upper bound, not what changed hands. Those observations are recorded but excluded from the accuracy maths, because pooling them with real transfer prices would bias every band in the flattering direction.
- Measurement may widen a range, never narrow one. Calibration is allowed to tell us we are more uncertain than we assumed. It is not allowed to license more precision than the assumption did. Tightening a band is a decision to take deliberately once hit rates are trusted, not a side effect.
- Stale scores stop counting. Buckets carry the date they were computed, so if the scoring job stops running the bands revert to the conservative default instead of quietly selling months-old figures as current measurement.
- A withdrawn listing teaches nothing about price. It is recorded as an outcome with no error, rather than guessed at.
What the confidence score is made of
Confidence is not a mood. It is computed from how directly the winning method measures your property multiplied by how much evidence stood behind it, then adjusted for what the report found:
- A like-for-like rate from genuine comparables, applied to your own floor area, is the strongest route. A raw suburb average is the weakest.
- Six comparables count for far more than one. A single comparable is capped hard, because one sale is an anecdote.
- Confirmed sales outweigh asking prices, and asking evidence is discounted toward what properties actually fetch before it can move the number.
- Agreement matters: comparables that cluster tightly earn more confidence than the same number scattered across a wide band.
- An independent verifier and a consistency judge can cap it. Two separate runs landing on the same number can raise it.
The consequence worth knowing: a low confidence score is not Guru being unhelpful. It is Guru telling you how hard to lean on the number.
Where Guru stops
A Guru report is an evidence-based estimate, not a formal, sworn or bank valuation, and it is not prepared by a registered property valuer. If you need a valuation for a bond application, an estate, a divorce or a court, you need a professional valuer and Guru is not a substitute for one.
There are also things no model can see. Nobody has been inside the property. A structural defect, an unapproved extension, a servitude, a difficult neighbour, a body corporate in trouble: none of that is in the data, and a confident number is not a survey. Guru reads the photos for condition and finish and flags what they suggest, but photographs are chosen by the seller.
And the market moves. An estimate is a read on the evidence available on the day it was produced. Read the AI disclaimer for the full picture.
What to do with all this
Use the range, not the midpoint. Read the confidence score before you read the number. Check which evidence route the report used, and if it says the data was thin, treat the estimate as a starting point for questions rather than an answer. Where Guru corrected itself, the report says so: that note is information, not an apology.
Then use it for what it is genuinely good at, which is knowing whether an asking price is defensible before you sit down opposite somebody who negotiates property for a living.
Frequently asked questions
Will you publish your error rate?
Is a Guru estimate a formal valuation?
Why did my range get wider after a re-run?
Why is the confidence lower than I expected?
More answers on the full FAQ page.