How a score is calculated
Release Health is Quash's score, out of 100, for the bugs that reach users in an app's recent builds. Quash builds QA automation, and computes this score from public Play Store reviews, with engine version 2.1.0. It is not the store rating, and it is not a live outage status.
The five dimensions
Two dimensions are required, because they carry 70 of the 100 points. If one of the three smaller dimensions cannot be measured, it drops out and the rest are rescaled. Each build is measured over its first 14 days, so a build that has been live longer is not punished for having collected more reviews. The window behind a score is the prior 180 days, and at most the last ten builds.
| Dimension | Points | What it measures |
|---|---|---|
| Severe defect rate | 35 | Share of reviews that report a severe defect, such as a crash, a freeze, or a payment that fails. |
| Regression severity | 35 | How much worse a build lands than the one before it, averaged over the builds that got worse. |
| Regression frequency | 10 | Share of releases that land at least 5 points worse than the one before. |
| Build consistency | 10 | The gap between the best and worst build in the window. |
| Defect persistence | 10 | The longest run of consecutive builds sitting above the app's own defect median. |
Why these weights
The weights come from a temporal holdout: each dimension was tested on how well it predicted problems in reviews it had not yet seen. Giving the three smaller dimensions more than 30 points between them made the score predict worse, and the drop was steady as their share grew. At 35, 35, 10, 10 and 10 the score predicted about as well as the two large dimensions on their own (a correlation of −0.552 against −0.562, and −0.694 against −0.699 on 92 apps), which is inside sampling noise. That is why the split sits there and not on rounder numbers.
The curves that turn a measurement into points were set on 7 September 2026 and frozen. A company's score does not move because its peers' scores moved. Changing a curve means re-running the holdout and publishing a new engine version.
Bands and ranking
- 80 and above is Battle-Tested.
- 60 to 79 is Solid Foundation.
- 40 to 59 is Bugs Slipping Through.
- Below 40 is QA at Risk.
An app with too few texted reviews, or fewer than three measured builds, is scored but left unranked. Ranking it would assert an order the data cannot support. It keeps its page.
What a page does not show
The specific defect themes, the versions they affect, the reviews behind them, and the tests that would have caught them are sent only after an email address is given. They are not in the page HTML. Hiding them with CSS would still leave them for a search engine to index.
iOS scores are not in this release. They are coming. Where an App Store listing is known, the page links to it. Android release dates are reconstructed from the versions reviewers were running, not from a changelog.
Updates and corrections
Scores are refreshed monthly, and every page states the date it was last scored. The roster of apps grows between runs.
If a score or a fact on your company's page is wrong, email prakhar@quashbugs.com with the app name and what should change. A person reviews the request before anything is republished.
Common questions
- What does a Release Health score measure?
- Whether recent builds are shipping bugs to users. The published number is built from public Play Store reviews and release history, not from the store's own star rating. Quash builds QA automation. An iOS listing is linked when we know it. iOS scores are coming.
- How often are scores updated?
- Monthly. Each page shows the date of its last scoring run.
- Why is an app scored but not ranked?
- A score needs enough reviews and enough measured builds to place one app above another. Below that, the app keeps its page and its score, and stays off the leaderboard.
- How do I correct a score or a fact?
- Email prakhar@quashbugs.com with the app name and what looks wrong. A correction is reviewed by a person before the page changes.