By the editors at Reliable
The short version: Among the sources reviewed for this guide, we found no U.S. government source that publishes a cost of bad data for manufacturing as a whole.
The closest federal figures are NIST studies of data exchange problems. One put the cost of imperfect design data exchange in the U.S. automotive supply chain at a minimum of $1 billion per year (1999). Another put inadequate supply chain data infrastructure at more than $5 billion per year in the automotive industry and almost $3.9 billion in electronics (2004).
The numbers quoted most often are cross-industry. The earliest version of the $3.1 trillion figure we found is a 2011 press release that presents it as one analyst’s estimate. IBM later repeated it in an infographic that has since been taken down. Gartner’s $12.9 million average comes from a 2020 survey in which customers of data quality software vendors estimated their own costs. The 15% to 25% of revenue range is a consultant’s synthesis of other estimates.
None of those three was measured in manufacturing plants.
The best-documented error rate comes from a peer-reviewed study of 75 data quality assessments. It found that 47% of newly created records had at least one critical error.
How We Evaluated
Independent editorial analysis based on publicly available government studies, peer-reviewed research, and the original pages behind widely quoted estimates.
Reliable Magazine does not sell data quality software or consulting and has no commercial interest in routing readers toward any particular vendor. Reliable does not accept payment for rankings. Vendors may sponsor enhanced listings with additional detail, but editorial rankings are independent. Read our editorial policy.
We sorted every figure into three tiers and kept the tiers apart:
- Government studies: NIST planning studies and economic briefs on manufacturing data exchange.
- Peer-reviewed research: one journal study of measured data quality.
- Analyst, vendor, and consultant estimates: IBM, Gartner, and Thomas Redman, with each party’s commercial stake stated at each use.
Three rules applied throughout. Each figure stays in the dollars its source reports, with no inflation adjustment. Nothing is summed across studies. A figure appears only if a named source page prints it.
The Closest Government Figures: NIST Data Exchange Studies
The National Institute of Standards and Technology (NIST) has commissioned several studies of what poor data exchange costs manufacturers. They measure interoperability and infrastructure gaps, which overlap with bad data without being the same thing.
Automotive design data (1999)
NIST’s 1999 planning study, prepared by Research Triangle Institute, examined product design data exchange in the U.S. automotive supply chain. It estimates that imperfect interoperability imposes a minimum of $1 billion per year on supply chain members.
The study abstract says the largest component of that cost, by far, is the resources devoted to repairing or reentering data files that downstream applications cannot use. It also calls the estimate conservative because some sources of cost could not be quantified.
For anyone tracing the phrase “bad data” to a federal number, this is the closest match: money spent fixing and rekeying files that arrived unusable.
Supply chain business data (2004)
A 2004 study prepared by RTI International for NIST examined the information that flows between supply chain partners in two industries. It puts the annual cost of inadequate supply chain integration infrastructure at more than $5 billion for the U.S. automotive industry and almost $3.9 billion for the U.S. electronics industry.
Smart manufacturing infrastructure (2016)
NIST Economic Analysis Brief 4 summarizes a study of technology infrastructure needs for smart manufacturing. It estimates that meeting those needs would save manufacturers $57.4 billion per year, about a 3.2% reduction in the shop floor cost of production. The brief calls the estimate conservative and reports it in 2013 dollars.
This is a projection of potential savings from better standards, measurement, and data infrastructure. It should not be quoted as a measured cost of bad data.
What these studies do and do not show
All three are government-commissioned and specific to manufacturing or manufacturing supply chains. All three are also dated, and each measures a defined slice of the problem: design file exchange, business transaction data, or missing technical infrastructure.
None of them measures the cost of wrong master data, bad sensor readings, or inaccurate work order history inside a single plant.
The $3.1 Trillion Figure
The most repeated number in this field is that bad data costs the United States $3.1 trillion per year.
The earliest version this guide found is a press release dated September 11, 2011. It attributes the figure to Hollis Tibbetts, then an analyst at Artemis Ventures, a software marketing and technology consulting firm. In the release, Tibbetts calls the number his estimate and allows that it could be off in either direction.
IBM later repeated the figure in an infographic. Thomas Redman cited IBM’s version in a September 2016 Harvard Business Review article, describing it as IBM’s estimate of the yearly cost of poor quality data in the U.S. in 2016. Redman is a data quality consultant, and IBM sells data management software.
Redman later described where he first saw the number. In a June 2023 note attached to a repost of the article, he says it appeared in an IBM infographic on the characteristics of big data about ten years earlier, and that IBM has since taken the infographic down.
He writes that he contacted IBM, that people there would not detail the methodology, and that they stood by the number. He also says he tested the figure against his own models and wrote the article partly to draw attention with a very large number.
By Redman’s account, $3.1 trillion equaled about 18% of the U.S. economy at the time.
Three points follow for a manufacturing reader. The figure covers the whole economy. The earliest version found is one analyst’s estimate with no published method. The IBM infographic that repeated it is no longer online.
The $12.9 Million Figure
Gartner states that poor data quality costs organizations at least $12.9 million a year on average, and its data quality topic page attributes the number to Gartner research from 2020. Gartner sells research and advisory services.
The figure comes from Gartner’s 2020 Magic Quadrant for Data Quality Solutions. A reprint of that report attributes the $12.9 million to Gartner’s customer reference survey, which included 154 organizations associated with 16 vendors.
That sample matters. The respondents are reference customers of the data quality vendors Gartner was evaluating, and the cost is their own estimate. Gartner does not state how many of the 154 organizations answered the cost question. The average describes that group across industries. The group is not a representative sample of organizations. Gartner does not present it as a manufacturing figure.
Gartner’s same page reports that 59% of organizations do not measure data quality, underscoring how difficult it is for many organizations to state their own cost precisely.
The 15% to 25% of Revenue Figure
In a 2017 MIT Sloan Management Review article, Redman estimates the cost of bad data at 15% to 25% of revenue for most companies.
The article explains how the range was built. It synthesizes three outside estimates: an Experian figure that bad data cost companies 23% of revenue worldwide, a per-employee cost estimate of $20,000, and an estimate that 16% to 32% of effort is wasted dealing with data.
Experian sells data quality products, and Redman consults on data quality. The range is an informed synthesis by a practitioner. It is not a measurement of any manufacturing sector, and the article does not claim to be one.
The 47% Error Rate
The best-documented evidence we found on how much data is wrong comes from a peer-reviewed study by Tadhg Nagle, Thomas Redman, and David Sammon, published in Business Horizons in 2020.
The study analyzed 75 data quality assessments collected over two years from a wide range of organizations, data sets, and business processes. On average, 47% of recently created data records had at least one critical error. Only 3% of the data quality scores were rated acceptable, defined as 97% or higher.
Redman’s MIT Sloan article describes the method behind those assessments: executives identified the last 100 units of work their departments had done, treated them as 100 data records, and reviewed the quality of that work.
The study reports an error rate and gives no dollar cost. Participants scored their own data, and the sample spans many sectors.
The 1-10-100 Rule
The 1-10-100 rule says it costs $1 to prevent a data error, $10 to correct it later, and $100 when the error causes a failure.
An Australian Government guidance page on data quality attributes the rule to the 1992 book Making Quality Work, written by George Labovitz, Yu Sang Chang, and Victor Rosansky. The rule began as a quality management rule of thumb and was later applied to data.
Treat the ratios as an illustration of how cost grows with delay. None of the pages reviewed for this guide cites a measured dataset behind them.
At a Glance
| Figure | Source | Year | What it measures | Manufacturing-specific | Provenance |
|---|---|---|---|---|---|
| $1 billion per year, minimum | NIST planning study, prepared by RTI | 1999 | Imperfect interoperability of product design data, U.S. automotive supply chain | Yes | Government-commissioned; described as conservative |
| More than $5 billion and almost $3.9 billion per year | NIST planning study, prepared by RTI | 2004 | Inadequate infrastructure for supply chain integration, automotive and electronics industries | Yes | Government-commissioned |
| $57.4 billion per year | NIST Economic Analysis Brief 4 | 2016 | Potential savings from meeting smart manufacturing infrastructure needs | Yes | Projection in 2013 dollars |
| $3.1 trillion per year | Tibbetts press release; later repeated by IBM and cited in Harvard Business Review | 2011 to 2016 | Cost of poor quality data to the U.S. economy | No | One analyst’s estimate; IBM infographic since removed |
| $12.9 million per year | Gartner | 2020 | Average cost of poor data quality per organization | No | Customer reference survey of 154 organizations across 16 vendors |
| 15% to 25% of revenue | Redman, MIT Sloan Management Review | 2017 | Cost of bad data for most companies | No | Synthesis of three outside estimates |
| 47% of records | Nagle, Redman, and Sammon, Business Horizons | 2020 | Newly created records with one or more critical errors | No | Peer-reviewed; 75 assessments |
Honest Limitations
- No direct measure exists in these sources. Nothing here supports a single cost of bad data for U.S. manufacturing, and this guide does not construct one.
- The NIST studies are dated. The interoperability figures come from 1999 and 2004, and this guide repeats them as reported. Software, standards, and supply chains have changed since.
- Interoperability and bad data overlap only partly. The NIST studies measure data that cannot be exchanged or reused. They do not measure data that moves cleanly and is simply wrong.
- The $57.4 billion figure is a projection. It estimates savings from infrastructure that did not yet exist when the study was done.
- The $3.1 trillion figure cannot be independently reproduced from a published methodology. The earliest version found is a 2011 analyst estimate, and the IBM infographic that repeated it has been removed.
- The Gartner sample is narrow. It reflects cost estimates from a reference survey of 154 customer organizations associated with 16 data quality vendors.
- Several figures share one author. Redman wrote or co-wrote three of the sources here and consults on data quality.
- The 47% study is self-scored. Managers reviewed their own records, and the organizations were not a random sample.
- Vendor surveys are left out. Recent surveys from data software companies report revenue percentages that this guide did not trace to a method.
Frequently Asked Questions
How much does bad data cost manufacturers?
No U.S. government source reviewed here gives a total. The closest federal figures are NIST studies of data exchange: a minimum of $1 billion per year for design data in the U.S. automotive supply chain (1999), and more than $5 billion in automotive and almost $3.9 billion in electronics for supply chain data infrastructure (2004). The widely quoted $3.1 trillion and $12.9 million figures are cross-industry and were not measured in manufacturing.
Where does the $3.1 trillion bad data figure come from?
The earliest version this guide found is a September 2011 press release that presents it as analyst Hollis Tibbetts’s own estimate. IBM later repeated it in an infographic, and Thomas Redman cited IBM’s version in a 2016 Harvard Business Review article. In a 2023 note, Redman says IBM has taken the infographic down, that IBM staff would not detail the methodology when he asked, and that they stood by the number. The figure describes the whole U.S. economy.
Is Gartner’s $12.9 million figure specific to manufacturing?
No. The figure comes from the customer reference survey behind Gartner’s 2020 Magic Quadrant for Data Quality Solutions, which included 154 organizations associated with 16 data quality vendors. The respondents were drawn from many industries, and the cost was their own estimate.
Does bad data really cost 15% to 25% of revenue?
That range is Thomas Redman’s estimate, published in MIT Sloan Management Review in 2017. The article says it synthesizes three outside estimates, including an Experian figure of 23% of revenue. It is a practitioner’s synthesis for most companies and was not measured in manufacturing.
What is the 1-10-100 rule?
It is a rule of thumb that preventing a data error costs $1, correcting it later costs $10, and letting it cause a failure costs $100. An Australian Government guidance page attributes it to the 1992 book Making Quality Work, written by George Labovitz, Yu Sang Chang, and Victor Rosansky. The ratios illustrate how cost grows with delay and are not drawn from a published dataset in the sources reviewed here.
How can a plant measure its own data quality?
The method behind the 47% study is simple to repeat. As Redman describes it in MIT Sloan Management Review, a team identifies the last 100 units of work its department completed, treats them as 100 data records, and reviews them for errors. The peer-reviewed study rated a score acceptable at 97% or higher, and only 3% of the 75 assessments reached that level.
Related Guides
- Best MES Platforms for Manufacturing
- Best ERP Systems for Manufacturing
- Best CMMS Software
- Best EAM Software
- How to Calculate OEE
Sources
- National Institute of Standards and Technology: The Economic Impact of Technology Infrastructure for Smart Manufacturing (NIST Economic Analysis Brief 4; savings estimate in 2013 dollars)
- RTI International: Economic Impact of Inadequate Infrastructure for Supply Chain Integration (2004 report prepared for NIST)
- Portland State University, PDXScholar: Interoperability Cost Analysis of the U.S. Automotive Supply Chain (abstract and record of the 1999 RTI report prepared for NIST)
- Newswire: Dirty Data Costs the US Economy $3.1 Trillion Yearly (September 11, 2011 press release with the Tibbetts estimate)
- Harvard Business Review: Bad Data Costs the U.S. $3 Trillion Per Year (2016 article citing the IBM estimate)
- SAP Community: Bad Data Costs the U.S. $3 Trillion Per Year (repost carrying Redman’s June 2023 note on the IBM infographic)
- Gartner: Data Quality: Best Practices for Accurate Insights (analyst firm page stating the $12.9 million average)
- IT Business Plus: Magic Quadrant for Data Quality Solutions (hosted reprint of the 2020 Gartner report with the survey description)
- MIT Sloan Management Review: Seizing Opportunity in Data Quality (2017 article with the revenue range and its inputs)
- Business Horizons, Elsevier: Assessing data quality: A managerial call to action (publisher page for the 2020 peer-reviewed study)
- Australian Government: Strengthen data quality (attribution of the 1-10-100 rule)









