← All Insights

Leaser architecture / Field note 003

Measuring what would not have happened.

Every intervention creates two futures. Operations reveal only one.

A shared path divides after a black intervention marker into a solid teal actual future and a transparent counterfactual future
The distance between the two futures is the effect.

The most important number in enterprise AI is the one reality never reveals.

After an AI system changes a campaign, follows up with a prospect, adjusts an incentive, or recommends a price, the business observes what happened next. It does not observe the parallel future in which the intervention never occurred.

That unseen future is the counterfactual. The difference between it and the result we can see is the incremental effect of the decision. Without a credible estimate of that difference, a closed learning loop can learn the wrong lesson with extraordinary efficiency.

What happened What would have happened =
What the decision changed

Attribution is not incrementality.

Attribution asks which touchpoint receives credit for an outcome. Causal measurement asks whether the outcome was changed by the intervention at all.

A signed lease can be attributed to a search campaign because search was the final recorded touch. That does not prove the campaign created the lease. The prospect may already have chosen the property after seeing a listing, visiting the neighborhood, reading reviews, or speaking with a friend. The campaign may have captured existing intent rather than created new demand.

Both questions matter. Attribution helps operators understand journeys and allocate credit. Incrementality determines whether an action produced value above the baseline. Confusing the two systematically rewards channels that appear near conversion and punishes the earlier forces that changed probability.

One operating history

Two plausible futures.

The counterfactual is not fictional. It is an estimate that must declare how it was constructed, what assumptions it depends on, and how uncertain it is.

A naïve loop learns from everything except cause.

Suppose a system increases media spend and signed leases rise. If it records the outcome as a win, it may learn to spend more. But leases might have risen because it was the first week of peak season, a competitor raised prices, new inventory came online, or the property simultaneously changed concessions.

The raw outcome bundles the intervention with every other force acting on the property. Feed that bundle back as a reward and the system learns correlation as policy.

This is particularly dangerous in autonomous systems. An error in a dashboard may mislead one quarterly review. An error in a reward signal becomes a repeated behavior. Closed-loop learning without causal discipline can make a system confidently wrong faster.

The causal maturity ladder

Say exactly what the evidence can support.

  1. ObservedThe outcome occurred after the action.

    Useful operational truth. No claim that the action caused it.

  2. AttributedThe outcome was connected to a journey or source.

    Credit is assigned under a declared model and window.

  3. ModeledA no-action baseline was estimated.

    Seasonality, trend, covariates, and comparable units inform the counterfactual.

  4. ValidatedThe effect survived a credible experiment.

    Randomized holdouts or strong quasi-experimental designs support an incremental claim.

Define the test before seeing the result.

A decision cannot be evaluated honestly if success is invented after the outcome arrives. Every meaningful intervention needs a measurement contract at the moment it is proposed.

Intervention

What exactly changes—and what stays fixed?

Unit

Property, market, audience, campaign, floorplan, or time window?

Outcome

Which canonical business metric decides success?

Window

When can an effect reasonably appear, mature, and expire?

Counterfactual

What represents the no-action future?

Uncertainty

What range of effects is consistent with the evidence?

Decision rule

What result changes policy, confidence, or autonomy?

Why multifamily is difficult

The signal is sparse. The world is noisy.

Long lag

Demand may take weeks to become a tour, application, and signed lease.

Small samples

Lease counts become thin when divided by property, floorplan, channel, and time.

Seasonality

Move cycles, school calendars, and local events can overwhelm a short comparison.

Interference

Prospects cross neighborhoods, devices, properties, and channels; treatment can spill over.

Concurrent action

Pricing, concessions, reputation, inventory, and media often change together.

Selection bias

Operators intervene where conditions are already unusual, making before-and-after comparisons deceptive.

Use the strongest design the decision can support.

Randomization remains the cleanest route to causal truth, but the unit of randomization must fit operations. In multifamily, that may mean matched properties or markets rather than individual prospects. It may mean rotating treatment across time, preserving permanent holdouts, or varying intensity instead of shutting a channel off.

When randomization is impractical, synthetic controls, matched comparisons, difference-in-differences, and time-series counterfactuals can estimate the missing future—provided their assumptions and pre-treatment fit are visible.

This is established measurement science, not AI theater. Google’s work on geo experiments shows how non-overlapping regions can become treatment and control units. Its work on Bayesian structural time-series models estimates the response that would have occurred without an intervention. More recent work on Meridian GeoX combines experimental design, multiple counterfactual methods, and placebo inference for real-world geo measurement.

The method should follow the decision—not the other way around. A high-risk, high-cost, irreversible action deserves stronger evidence than a cheap, reversible adjustment. When the expected effect is smaller than the noise the design can detect, the honest answer is not “no impact.” It is “not measured precisely enough.”

The prerequisite

Causal inference begins with outcome truth.

No statistical technique can rescue a mislabeled outcome. A platform conversion is not a signed lease. A lead is not occupancy. A model is only as honest as the entity, identity, timing, and metric definitions beneath it.

ImpressionVisitLeadTourApplicationSigned lease

Leaser’s attribution spine resolves demand activity to the canonical property, floorplan, prospect, and lease. A shared semantic layer keeps cost per signed lease, exposure, occupancy, and NOI consistent wherever they appear.

Saving spend can be causal value.

Consider a property whose demand is pacing ahead while one paid channel continues to spend. Leaser proposes a bounded reduction. During the measurement window, signed-lease velocity remains stable.

A before-and-after view says “nothing changed.” A causal view asks whether velocity would have risen, fallen, or stayed flat without the cut. A matched no-action baseline indicates the same leasing trajectory would probably have continued. The intervention therefore created value by preserving the outcome with less spend.

Now reverse the result. If the treated property falls below its counterfactual by more than the savings justify, the action was destructive even though cost per lead improved. The metric that looked efficient locally reduced the outcome the system exists to produce.

The intelligent loop

Learn from lift, not applause.

An operator accepting a recommendation is not evidence that it worked. A metric moving afterward is not enough either. The learning signal should compare expected impact, observed outcome, and estimated no-action outcome.

  1. ExpectedWhat did the decision predict?
  2. ObservedWhat did the business record?
  3. CounterfactualWhat likely happens without action?
  4. IncrementalWhat changed because we acted?
  5. LearnedWhat should the next decision believe?

Measurement maturity must be visible.

Leaser already connects leads to canonical lease outcomes with explicit confidence, centralizes the metric definitions that operators and models share, and stores daily occupancy inputs so forecasts can be replayed against what later occurred.

We are equally explicit about what is not yet causal. Today’s attribution stitch is single-touch: different model labels do not yet produce different matching behavior. Multi-touch journeys, channel-role-weighted contribution, real counterfactual monitoring, and geo incrementality are advancing capabilities—not claims we disguise as finished truth.

That candor is architectural. Leaser’s evidence system separates observed, estimated, modeled, and validated claims. Coverage, sample size, freshness, confidence, and measurement method travel with the decision. The interface should never make a correlation look experimentally proven.

Not every decision needs an experiment.

Measurement itself has a cost. Holding out demand can sacrifice near-term performance; waiting for statistical power can delay action; over-segmenting can make every result inconclusive.

The goal is not to turn every operating choice into a research program. It is to match rigor to consequence. Low-risk reversible actions can learn from repeated modeled comparisons. Portfolio-wide policies and autonomy thresholds need stronger validation. Experiments should be concentrated where uncertainty is expensive and the result can improve many future decisions.

The compounding asset is not a history of actions.

It is a memory of interventions, counterfactuals, and effects.

Over time, Leaser can learn not simply that a play was followed by a lease, but where the play created incremental value, for which property conditions, with what delay, at what confidence, and within which constraints. That memory transfers cautiously across a portfolio and improves the design of the next test.

Outcome Intelligence demands more than seeing what happened.

It must measure the future we changed.

Build the measurement layer with us