Methodology
The Hedonic Factor Model
At each point in time we run a cross-sectional OLS regression of log(price/sqm) on property characteristics using a rolling 3-month window. The fitted coefficients (betas) tell us the market price of each characteristic in that window.
For each rolling window t and each transaction i, the regression is:
log(pi,t / ai,t) = αt + Σk βk,t · xk,i,t + εi,t
where p is the transacted price, a the floor area in square metres and xk the k-th factor (continuous or dummy). Every factor is centred on its average over the whole period, so the intercept αt is the log price per sqm of the average home.
Vocabulary
Every page and post uses these terms, and only these:
- Exposure: how much of a factor a home has (xk).
- Premium: the % price-per-sqm gap between two otherwise identical homes that differ by one stated unit of a factor: exp(βk,t · unit) − 1. Each factor card states its unit, for example a house against a flat of the same size, or twice the floor area.
- Factor Return: the % change in the Premium since the factor's first full 3-month window. It is the return on holding one fixed unit of the factor, so it moves only when the market reprices the factor. (Factors page)
- Index and Index Return: a constant-quality price index and its % change. (Trends page)
- Baseline Market: the part of the Index Return that no factor explains (the change in α).
- Factor Contribution: a factor's part of the Index Return.
- Average Price Paid: the average price per sqm of the homes actually sold (geometric mean), which moves when the mix of homes sold changes.
- Composition: a factor's part of the gap between the Average Price Paid and the Index, from changes in what sold.
The Index and its decomposition
The Index reprices a slowly updated basket of homes, x̃t: a blend of the average home sold over the past 24 and 6 months, known one month ahead. Chain-linking month by month:
Δ log Indext = Δαt + Σk x̃k,t · Δβk,t
Index Return = Baseline Market + Σk Factor Contributionk
Average Price Paid = Index Return + Σk Compositionk
Both identities hold exactly. The log parts are rescaled so that they also add up in %, and every series is measured from the first full 3-month window. Because the basket changes slowly, a burst of sales of one kind of home moves the Average Price Paid but not the Index; that difference is the Composition.
- Factor Returns do not add up. Each is the change in one factor's Premium at its own unit; quote each independently. Factor Contributions and Compositions do add up, through the identities above.
- Period changes are not simple differences of cumulative percentages. Use (1 + latest/100) / (1 + prev/100) − 1, not latest − prev.
Why a rolling 3-month window
- Sample size. A single calendar month doesn't give enough transactions per postcode × type cell to fit a stable regression. Three months roughly triples the cross-sectional sample without forcing us to look at quarterly data.
- Seasonality. Monthly volumes are seasonal but the 3-month rolling cut largely smooths out the dip-and-bounce pattern.
- Stability. Coefficients on rarer factor cells (EPC A/B, new build, certain age bands) move violently with single-month windows. Three months keeps the marginal-buyer signal visible without exploding the variance.
Factor Selection
Factors are selected iteratively using forward selection. At each step, candidate factors are evaluated over all rolling windows. A factor is accepted if:
- Median p-value across all windows < 0.10
- Significant (p < 0.10) in ≥ 50% of windows
- Correlation of its month-on-month contribution with each already-accepted factor's < 0.50
- Its exposure is not collinear with the accepted factors: pairwise correlation ≤ 0.80 and every variance inflation factor ≤ 5, so no two factors take turns explaining the same thing
- It must not be built from the sale's own price. A former "price tier" factor bucketed each sale by its own price per sqm, which restates the quantity being explained; it was removed from every city in September 2026.
This ensures each factor adds independent, stable information to the model.
Per-city feature engineering
The set of hedonic characteristics differs by city because the underlying datasets differ. Where multiple encodings are plausible we prefer the simplest one that produces an interpretable coefficient.
- London. Each sale joined to its own EPC certificate by address. District price level (location), outer vs inner London, size (log floor area), house vs a flat of the same size (which carries tenure: 99% of freeholds are houses), construction era (U-shape), energy rating (learnt only from certificates within two years of the sale) and recent build (planning data matched by location).
- New York. Neighbourhood price level, size (log floor area), construction era (U-shape), elevator building, building height, mid-rise apartment building and recent construction (2010+).
- Paris. Arrondissement price level, floor area (linear, plus distance from 60 sqm and a large-home flag), room size (sqm per room), construction era (U-shape) and the GES emissions label.
- Singapore. HDB resale and URA private together: the private-vs-public segment, building age (log and distance from the median), floor (distance from the median and a low-floor flag), size (distance from 60 sqm, plus studio, mid-size and large flags) and the 2000-09 lease cohort.
- Taipei. District price level, floor area, apartment vs house, walk-up and mid-rise mansion building types, and recent construction (2015+).
- Seoul. District (gu) price level, size (log and distance from 70 sqm), floor level, building age, construction era (U-shape) and pre-1990 redevelopment candidates.
- Tokyo. Ward price level, building age, station distance (linear and a 15-minute-plus flag), size (log plus a large-unit flag), room layout, steel-reinforced concrete structure and recent construction (2015+).
Quality Controls
- Top 0.5% by price/sqm removed each window (data errors and ultra-prime outliers)
- Observations with z-score < −5 on log(ppsqm) removed (suspiciously cheap transactions)
- Minimum 50 transactions per window; sparser windows are skipped
- Non-linear encodings (U-shapes, collapsings) impose economic priors to prevent fitting noise
- Standard errors and p-values are stored alongside the coefficients and used in factor selection
Limitations
- The model is unconditional on macro variables (rates, GDP, etc.). We measure the moves; we don't attribute them to a specific cause beyond what the hedonic decomposition gives us.
- The most recent 2 to 3 months of each series should be treated as provisional: the rolling window hasn't fully updated and late-arriving registrations will revise figures slightly.
- Outside the cities we currently cover, the same methodology can be applied wherever a transaction-level dataset with property characteristics is publicly available.
References
The hedonic approach to residential property pricing goes back to Rosen (1974). The classic survey is Sirmans, Macpherson and Zietz (2005), The Composition of Hedonic Pricing Models, Journal of Real Estate Literature 13(1).