Software Bug & Defect Statistics (2026)

Nishtha chauhan
Nishtha chauhan
|Published on |7 Mins
Cover Image for Software Bug & Defect Statistics (2026)

A release goes live, a customer reports a failure, and someone asks the question that sounds simple: how many bugs does software actually have? The available numbers rarely settle it. They may count static-analysis findings, confirmed defects, outages, maintenance effort, or modeled economic loss—different measurements with different denominators.

The short answer is that there is no verified universal software-bug rate for 2026, nor a universal percentage of defects that reach production. You can, however, use the statistics below when you preserve what each one measures: its population, period, and evidence type. That makes the figures useful in a budget case, a QA metric discussion, or a cited report without turning them into an invented industry average.

Key software bug statistics for 2026

  • About 2,100 reliability issues per million lines of code: Sonar observed this rate in code analyzed during the last six months of 2024, covering more than 7.9 billion lines of code, over 970,000 developers, and 40,000+ organizations. It is a first-party static-analysis observation among SonarQube users, not a universal functional-bug rate. Sonar’s report names seven languages in scope.

  • At least $2.41 trillion: CISQ’s modeled estimate for the U.S. cost of poor software quality in 2022. CISQ separately estimated accumulated software technical debt at about $1.52 trillion. These are modeled aggregate economic estimates, not a current cost per defect. CISQ.

  • $59.5 billion annually: a historical estimate from a 2002 study prepared for NIST of the U.S. cost of software flaws. NIST later summarized the study and its finding that testing consumed 50% of software-development budgets. This is not an inflation-adjusted 2026 loss. NIST.

  • 44% of systems below a recommended maintainability threshold: SIG reports this for the systems in its maintainability model, alongside unnecessary maintenance costs of up to EUR250,000 per system per year and up to EUR7 million annually for large enterprise systems. It is a code-maintainability benchmark, not a defect-incidence rate. SIG.

  • 2,417 Maintenance & Support applications: ISBSG says its 2026 data release includes this many applications, alongside 13,718 Development & Enhancement projects. Its repository includes defect-count fields, but its public page does not publish an aggregate defect rate. ISBSG.

  • More than 14,000 defects: a 2026 academic study examined this many defects across open-source C/C++ and Java systems, screening more than 200,000 candidates and identifying about 7,000 confirmed residual defects. The study is evidence about characteristics of post-release defects, not the share of all defects that escape testing. Cotroneo et al..

  • Three years of production-incident data: a 2025 Google Research study examined outages in planet-scale big-data and machine-learning services serving billions of users. An outage is an operational reliability event, not a count of individual code defects. Google Research.

Ebook Preview

Get the Mobile Testing Playbook Used by 800+ QA Teams

Discover 50+ battle-tested strategies to catch critical bugs before production and ship 5-star apps faster.

100% Free. No spam. Unsubscribe anytime.

Software bug statistics at a glance

Metric

Figure

Population and period

Evidence type

What it does not measure

Reliability issues

About 2,100 per million lines of code

SonarQube-user code analyzed in the last six months of 2024

First-party static-analysis observation

A universal bug rate

Cost of poor software quality

At least $2.41 trillion

U.S. estimate, 2022

CISQ modeled aggregate estimate

Cost per individual defect

Accumulated technical debt

About $1.52 trillion

U.S. estimate, 2022

CISQ modeled aggregate estimate

Defects in production

Software-flaw cost

$59.5 billion annually

U.S. study, 2002

Historical estimate summarized by NIST

Current-dollar 2026 cost

Maintainability threshold

44% below threshold

Systems assessed in SIG’s model, reported 2025

First-party benchmark/model

Defect incidence

Maintenance-support repository

2,417 applications

ISBSG 2026 release

First-party repository description

An aggregate defect rate

Post-release defect study

More than 14,000 defects; about 7,000 confirmed residual defects

Open-source C/C++ and Java systems, 2026 paper

Primary academic research

A universal escape percentage

Production outages

Three years of incident statistics

Planet-scale services, 2025 paper

Primary research

Individual bug counts

How many bugs does software contain?

No credible source in this evidence set establishes one number for all software. A bug count changes with the definition of a bug, the detector, the language, the age of the system, the release process, and the denominator used. Defects per thousand lines of code, defects per story point, findings per repository, and customer-reported incidents can each be valid internal measures without being comparable to one another.

The closest large-scale current benchmark here is Sonar’s observation of about 2,100 reliability issues per million lines of code. Sonar’s population was substantial—more than 7.9 billion lines of code across more than 970,000 developers and 40,000+ organizations in the last half of 2024—but it remains code analyzed by SonarQube and findings identified by its rules. Sonar does not present the result as a census of all production defects.

If you need a defensible sentence for a deck, use the denominator and population: Sonar observed about 2,100 reliability issues per million lines of code in SonarQube-analyzed code during the last six months of 2024. That is narrower than “software contains 2,100 bugs per million lines,” and it is much more useful.

What software security statistics can—and cannot—tell you

Security flaws and functional defects overlap, but they are not the same metric family. A security scan may flag a vulnerability pattern that has no visible user-facing failure; a functional defect may have no security consequence. Combining both into a single software-bug rate makes neither number clearer.

This report does not publish a 2025 application-security prevalence percentage because the available primary material does not expose a precise figure with its necessary scan type, application population, and measurement period. The absence of a citable percentage is preferable to repeating a headline number without its denominator.

For your own reporting, keep security findings as a separate trend: define the scanner and ruleset, severity threshold, application population, and time window. That lets you determine whether your security exposure is changing without claiming it is the same thing as escaped functional defects.

What do software bugs cost?

A cost statistic is often used as shorthand for the impact of bugs. The published figures here describe different economic questions, so they should stay separate.

CISQ’s modeled U.S. estimate

CISQ estimated the 2022 U.S. cost of poor software quality at at least $2.41 trillion and accumulated technical debt at about $1.52 trillion. CISQ’s 2022 report labels these as broad estimates of poor quality and technical debt. They are not a count of defects, and dividing them by an assumed number of bugs would manufacture a per-bug cost that CISQ did not publish.

NIST’s historical software-flaw estimate

NIST summarizes a 2002 study prepared for the agency that estimated annual U.S. software-flaw cost at $59.5 billion. NIST’s 2010 account is valuable historical context, but the figure is over two decades old. Do not present it as a 2026 market loss or quietly convert it into a current estimate.

SIG’s maintainability benchmark

SIG reports that 44% of systems fall below its recommended maintainability threshold. It also says unnecessary maintenance costs can reach EUR250,000 per system annually, and as much as EUR7 million annually for large enterprise systems. SIG is describing costs associated with its maintainability model. Use it to discuss ongoing code-quality burden, not to estimate how many defects a system contains.

The practical takeaway is to connect defect data to your own impact data: support contacts, affected users, recovery time, rework hours, lost transactions, or contractual exposure. A raw bug count does not tell you which failures matter most.

What the evidence says about bugs reaching production

There is no verified universal percentage for defects that escape into production. That missing number is not a gap you can solve by averaging unrelated studies: their populations, definitions, and release windows differ too much.

The 2026 study by Cotroneo and coauthors provides a more useful, bounded result. Across open-source C/C++ and Java systems, it examined more than 14,000 defects and identified about 7,000 confirmed residual defects after screening more than 200,000 candidates. The researchers found that post-release defects were associated with older, frequently modified, high-churn components, and that their fixes were typically longer and more complex. The study helps you prioritize riskier components; it does not say what percentage of defects escapes in every organization.

Google Research offers a second, operational lens through a 2025 study of three years of production outages in planet-scale services. Its paper treats outage patterns as software-reliability data. Outages can reflect defects, deployment conditions, detection behavior, and operational dependencies, so they are not interchangeable with bug reports.

For your releases, define an escaped defect before you count it. For example, you might count confirmed defects first detected by customers within 30 days of release, then divide by all confirmed defects discovered for that release. The definition is yours to set, but it must remain stable if you want the trend to mean anything.

Why software bug statistics disagree

The disagreement usually starts before any arithmetic: the sources are counting different events.

  • Units differ. A reliability finding per million lines of code, a defect ticket, an affected application, an outage, and a dollar estimate are not interchangeable units.

  • Detectors differ. Static analyzers vary in their coverage and false-positive behavior. NIST’s SATE V evaluation examined 14 static analyzers and found considerable differences in bug-class support, findings, and wrong-finding frequency. NIST.

  • Populations differ. SonarQube users, SIG-assessed systems, ISBSG contributors, open-source C/C++ and Java projects, and planet-scale services are not samples of the same universe.

  • Event definitions differ. A static-analysis finding may be actionable code debt; a residual defect is a defect that reached a post-release context; an outage is a reliability event. Each calls for a different numerator.

  • Periods differ. This report includes observations published from 2025 to 2026, source-code data from late 2024, a 2022 modeled cost estimate, and a 2002 historical estimate. Recency does not erase what a metric was designed to measure.

No first-party Quash telemetry, recurring bug-pattern data, customer language, or completed experiment currently establishes a universal bug rate or production-escape rate. The public record would need a controlled multi-application census with defined detectors, populations, and release windows to support that kind of claim.

What you should measure instead of one bug rate

Your QA metrics should answer operational questions, not chase an industry number that does not exist. Start with a small set of definitions your team can apply consistently.

  1. Defect density: Choose a denominator that fits your delivery model—such as lines of code, completed work items, or active users—and document it.

  2. Escaped-defect rate: Set a release window and count confirmed defects first found after release. Keep it distinct from deployment failures.

  3. Time to detect and time to repair: Track when an issue was introduced, detected, triaged, fixed, and verified where those timestamps are available.

  4. Severity and impact: Record affected users, revenue exposure, blocked workflows, and data or security implications alongside severity.

  5. Regression and reopen rates: These show whether fixes remain fixed and whether your validation process catches recurrence.

  6. Component and environment context: Capture the service, device or platform, language, release, and test environment associated with the issue.

The quality of the input determines the quality of the metric. A consistent, reproducible report makes trend analysis possible; this guide to writing effective bug reports explains the fields that make issues easier to reproduce and investigate.

Methodology and source limitations

This report separates static-analysis findings, modeled economic estimates, maintainability benchmarks, repository descriptions, post-release defect research, and production-incident research. It does not convert any of them into a universal rate, because their populations and event definitions do not support that conversion.

It also excludes a precise current security-prevalence statistic rather than repeating one without an extractable primary passage showing its measurement shape. That is a limitation of the public evidence, not evidence that security findings are rare or unimportant.

For a number to be comparable in your organization, record the metric, numerator, denominator, population, period, detector or reporting route, and inclusion rules. Change any of those fields and you may have a new metric rather than a trend.

Conclusion

Software bug statistics are decision metrics, not universal constants. Sonar’s 2024 reliability-finding rate can anchor a conversation about analyzed code; CISQ, NIST, and SIG illuminate distinct cost and maintainability questions; Cotroneo and Google Research provide bounded views of post-release defects and production reliability.

Use the figure that answers your actual question, preserve its population and period, and measure your own release outcomes with stable definitions. That gives you a statistic your organization can act on instead of a headline it cannot trust.