How the Greater Toronto Area report works.
This is not a critic’s list and it is not a star-rating clone. It is a structured reading of 409,349 public reviews across 319 GTA shawarma locations.
Reviews become evidence, not votes.
Written reviews are classified into food quality, meat, sauce, bread and freshness, portion, affordability, value, and consistency. A review only influences a component when it actually discusses that component. At most one contribution per review per component, so one long review cannot outweigh ten short ones.
Small samples are not allowed to cosplay as certainty.
Every component is Bayesian-smoothed toward the GTA-wide prior. Evidence counts and confidence labels remain visible. Locations with enough total reviews receive an official rank; smaller samples remain provisional, and locations without enough written review evidence are shown as unscored rather than assigned an invented score.
Every location gets exactly one category.
Locations are classified as shawarma-primary, shawarma-on-menu, or novelty/hybrid. Shawarma-primary locations can draw on their whole review corpus; shawarma-on-menu and novelty/hybrid locations are only scored from shawarma-relevant evidence within their reviews.
Coverage and confidence are disclosed, not hidden.
Each location shows the ratio of downloaded to expected reviews and a coverage status (complete, slightly short, materially incomplete, or zero). Restaurants that could not be safely resolved to a current Google listing, or that closed, were excluded rather than force-matched — see the 0 entries in this market’s reconciliation record.
The Toumometer is a joke with a methodology.
Only explicit evaluative mentions of toum or garlic sauce count. Those mentions are recency-weighted and smoothed, which is far too much statistical ceremony for a garlic-sauce leaderboard. That is the point.
Model version and limitations.
This market was scored with a compatible reimplementation of the published LSR-review-score-v1.0 methodology, recovered from the London workbook’s formulas, weights, and thresholds. The private London sentence-model code was unavailable, so a transparent lexical text score on the same 0–10 scale stands in for the 85% text component.
Reviews were analyzed in English; non-English reviews were retained in the corpus but are not yet scored by a validated language-specific model. This affects Montréal most, where a meaningful share of the corpus is French.