Skip to content
All articles

CRM & Leads

Lead Scoring for Real Estate: How to Identify Your Best Leads

Which leads deserve attention first: the signals worth weighing, the ones that mislead, why scores decay, and why its most-cited statistic doesn’t hold up.

By 12 min read
Lead Scoring for Real Estate: How to Identify Your Best Leads

Introduction

You have more leads than hours. That is the entire problem, and no amount of software makes it go away — it only changes which twelve people you call today.

Lead scoring is a way of making that choice deliberately instead of by whoever emailed most recently. It is worth understanding properly, and it is also worth being clear about what it is not: a score is a suggestion about ordering, not a judgement about people.

A note on this article. This is an educational piece. It describes how scoring works as a discipline. RealFoyer scores lead health, across budget, activity, engagement, response and timeline fit. It also records the behavioural signals described below — properties viewed, saved searches, listing alerts, enquiries — on the lead record, where a person can read them alongside the score. Where that distinction matters below, it is stated.

⚡ Quick answer — what should a lead score be built from? Four families of signal, weighed together: fit (can they transact), intent (what have they told you they want), engagement (what are they actually doing), and recency (is any of it still true). Behaviour without fit produces enthusiastic browsers. Fit without behaviour produces people who aren't moving. And without decay, every score drifts upward until the list is meaningless.

Why score at all

If you have thirty leads a month, you do not need scoring. You need a list and a habit. Scoring earns its place when volume exceeds the attention available — when the honest answer to "who should I call now" is "I don't know, and I'll pick the one I remember."

What scoring is really doing is making an implicit ranking explicit. You are already ranking. The question is whether the ranking is examinable.

That framing also sets the limit. A score should never be the only thing that decides whether someone is worth talking to. It should decide the order of a list a person still reads.

The signals

Fit

Can they transact?

  • Ownership tenure
  • Equity position
  • Stated budget vs inventory
  • Location served

Slow-moving and easy to over-weight — fit alone never means ready

Intent

What have they said they want?

  • Saved search created
  • Saved search modified
  • Valuation request
  • Direct enquiry

A modified search is stronger than a created one, and usually untracked

Engagement

How are they behaving?

  • Properties viewed
  • Repeat views of one listing
  • Listing alerts opened
  • Reply to an agent

Browsing is cheap; volume is not intent

Recency

Is this still true?

  • Days since last session
  • Return after silence
  • Time since last contact

Without decay, every score drifts upward forever

Priority

Who to call first today — an ordering, not a verdict, and one a person should be able to overrule without explaining themselves

An industry framework, shown for explanation — a way to think about the problem rather than a description of any one product. RealFoyer scores lead health from budget, activity, engagement, response and timeline fit.

An educational framework, not a product workflow. Four families weighed together — none of them sufficient alone.

Engagement — what they're doing

RealFoyer lead record with a lead health panel breaking a score into budget, activity, engagement, response and timeline components
A score you cannot decompose is a score nobody trusts. Each component here is a signal the section describes, scored separately. Demo data, contact details redacted.
SignalWhat it indicatesHow it misleads
Properties viewedActive considerationBrowsing is cheap; volume is not intent
Repeat views of one listingReal interest in a specific homeCan be one person showing a partner
Saved search createdCriteria formingOften set once and abandoned
Saved search modifiedCriteria firmingRarely tracked, and one of the strongest available
Listing alert openedSustained attentionOpen tracking is increasingly unreliable
Reply to an agentWillingness to engageStrong, and rare enough to weight heavily
Return visit after silenceRe-entry into the marketStrong, easy to miss entirely

The two rows worth pulling out are the ones most systems ignore. A modified saved search means someone reconsidered their own criteria — that is a person actively working the problem. A return after silence means a dormant record just became live, and it will not appear on any list sorted by lead age.

Intent — what they've told you

A saved search is a written statement of intent that the person composed themselves. A valuation request is a seller signalling before they are ready to say so. A direct enquiry about a specific listing is the strongest and least ambiguous signal available, and it is the one that most often sits in a queue.

Fit — whether they can transact

Behaviour answers "are they engaged." It does not answer "can they move."

Someone who has viewed forty properties may be six months from a mortgage approval. Someone who viewed two may have owned their home for eleven years and be sitting on substantial equity. Ownership tenure, equity position, and whether their stated budget matches anything in your inventory are all slow-moving and all easy to under-weight because they are less exciting than a live session.

One caution on fit signals: score what people tell you about themselves, in their own words, recorded in your notes. Do not infer household composition, family status, or anything adjacent to a protected characteristic. That is a fair housing exposure, and it is not a technicality.

Recency — whether it's still true

Scores must decay. A lead who was extremely active in March and silent since is not a hot lead in August; they are a record with a history.

Without decay, every score only ever goes up, everyone eventually crosses the threshold, and the list stops discriminating. The decay rate is a judgement call — residential cycles run long, so aggressive decay is as wrong as none — but the principle is not optional.

What should subtract

Most scoring models only add. A few negative signals are worth encoding:

  • A bad phone number or a bounced email. Not a judgement about the person, just a fact about reachability.
  • An explicit "not yet." If someone said they're looking next spring, honour it. Calling them in October because a number went up is how you lose them.
  • Opted out of contact. This should suppress, not deprioritise.
  • Duplicate of an existing active record. Merge rather than double-count.

Manual, rules-based, or predictive

Three approaches, in ascending order of what they demand from you.

Manual. A person reads the record and decides. Under a few dozen leads a month this is not a fallback — it is the best available option, because a person notices things no model encodes.

Rules-based. You write the weights. Valuation request +30. Saved search modified +20. No session in 60 days −15. Transparent, debuggable, and you can explain to an agent why a lead surfaced. Most teams should stop here.

Predictive. A model learns weights from your own closed deals. This needs data you probably do not have. Fitting a model to a small number of closings mostly learns the idiosyncrasies of those specific transactions, and it does so with enough confidence to be convincing. There is also a cold-start problem: a new office has no closings to learn from, and buying a generic model trained on someone else's market imports their assumptions along with their weights.

This is where AI scoring conceptually fits — and it is worth naming the requirement plainly. A predictive model needs volume, clean history, and outcomes recorded honestly. If "lost" is never marked in your CRM, there is no training signal, and a model trained on that data learns that nothing ever fails.

Setting the threshold

Derive the threshold from capacity, not from the score distribution.

Work out how many meaningful conversations your team can have in a week. That number, not a round score, is the threshold. If you can make forty calls, the threshold is wherever the fortieth lead sits.

Setting it any other way produces one of two failures: a "hot" list of two hundred that nobody works, or a list of four that leaves capacity idle.

Why scoring fails in practice

Rarely because the maths is wrong.

Nobody trusts it. One bad surfaced lead and agents revert to their own instincts permanently. Build in an override and make it visible — a score that can be argued with survives; one that can't gets ignored.

It scores what's easy to capture. Email opens are easy and nearly meaningless. A note saying "her lease ends in April" is hard to capture and worth more than any behavioural signal on the record.

It never gets checked. A model set up once and never validated against outcomes is a random number generator with a reputation.

It replaces judgement instead of ordering it. The score decides who to call first. It should not decide who is worth calling.

The statistic you should stop repeating

If you have read anything about lead response, you have met this figure: 391% higher conversion when responding within one minute, credited to an MIT study.

It is worth knowing what that citation actually is. The MIT and InsideSales.com Lead Response Management study is real and retrievable. The 391% figure is Velocify platform data — a different source. The most-repeated number in this discipline is routinely attributed to an institution that did not produce it.

The genuine study also carries a limitation almost nobody repeats: its dataset skewed toward mortgage, insurance and education — high-volume, phone-heavy verticals where contacting a lead within minutes is standard practice. Residential resale is not that. The transaction is rarer, larger, and far less phone-driven. A finding from a high-frequency phone vertical may or may not transfer to one where a client transacts twice a decade, and the study does not claim it does.

Then there is the orphan family — figures in wide circulation with no traceable source:

FigureUsually attributed toTraces to
"78% buy from the first responder"McKinsey, or InsideSales, or ForresterNone of them, verifiably
"35–50% of sales go to the first responder"The same three, interchangeablyNone of them, verifiably
"Respond within 5 minutes → 3–5× conversion"UnattributedNothing
"30% of CRM data is wasted"UnattributedNothing

Three different attributions for one number is the signature of an orphan statistic. A real finding has one source. Variation in the citation is people guessing at a reference for something they know they read somewhere.

None of this means responding quickly does not matter — it plainly does, for reasons needing no statistic. Speed to lead covers the mechanism.

Where RealFoyer sits in this

Worth being exact, because this is precisely the kind of article where a vague product reference becomes an accidental claim.

RealFoyer automation builder with a scoring step feeding a threshold block, and an execution preview that runs against a sample lead without sending anything
The dry-run preview matters more than the scoring step. A threshold you have not tested against real records is a guess with a number attached.

RealFoyer records a number of the signals above on the lead record: properties viewed, saved searches, listing alerts, enquiries, and the activity timeline that puts them in order. An agent can read all of it before making a call.

RealFoyer does score lead health, built from budget, activity, engagement, response and timeline fit, and surfaces it on the lead record as a priority signal. The families set out above are a wider way of thinking about the problem than any one implementation covers, so the reading still matters: a score orders your day, and a person decides what to do with the order.

Your first thirty days

  1. Write down how you currently decide who to call. You have a heuristic. Make it explicit — it is your version one.
  2. Inventory what your system actually captures. Not what it could capture. What is on the record today.
  3. Pick five signals. Two engagement, one intent, one fit, one recency. Five is enough.
  4. Set the threshold from capacity. How many real conversations per week?
  5. Run it manually for a month before automating anything.
  6. Check it against outcomes. Of the leads you prioritised, how many became conversations? If the answer is no better than random, the weights are wrong.

FAQ

How many signals should a score use? Five to eight. Beyond that you cannot explain why a lead surfaced, and unexplainable scores get ignored.

Should buyers and sellers use the same model? No. Seller signals — valuation requests, ownership tenure, equity — are structurally different from buyer browsing behaviour. Two models, or one with a type branch.

Can I score leads without website tracking? Partially. You lose engagement entirely and keep fit and stated intent. That is a weaker but still useful ordering.

Is a low score a reason not to call? No. It is a reason to call them fourth. A score orders a list; it does not remove anyone from it.

What about leads from portals with no behavioural data? Score them on fit and on what the enquiry itself says. And note that an enquiry sent to several agents at once has a reachability window, which is an argument about sequence rather than about score.

How often should the model be reviewed? Quarterly, against closings. If prioritised leads convert no better than the rest, the model is decoration.

In short

Scoring is a way to make an ordering you are already doing examinable. Build it from fit, intent, engagement and recency; let it decay; derive the threshold from capacity; and check it against outcomes or stop pretending it works.

And treat the number as an ordering of a list a person still reads — not as a verdict about who deserves a phone call.

Research Integrity

Rewritten 16 August 2026. This article is educational. The dedicated section states plainly what RealFoyer scores — lead health across budget, activity, engagement, response and timeline fit — rather than leaving it to inference, and the signal-model diagram is captioned as an industry framework rather than a description of that implementation. The 391% response-time statistic is retained only as a correction: the MIT/InsideSales Lead Response Management study is genuine and retrievable, the 391% figure originates with Velocify, and the study's dataset skews toward mortgage, insurance and education rather than residential resale. Four further figures in wide circulation are listed as untraceable rather than repeated as fact. Fifteen illustration specifications and one screenshot specification that were being published as article body text have been removed. No product CTA appears, and none will be added: a product link on a scoring article would imply a capability that does not exist. Related reading only.