Our Cost & Data Methodology
How Abode Sources, Verifies, and Updates National, Regional, and Local Home Improvement Cost Figures & Data
How we source our cost information
Pricing in home services varies enormously. The same job can differ by hundreds or thousands of dollars depending on your region, the age and condition of your home, the materials you choose, the season, and which contractor you hire. A range that holds in Cleveland can be well off in San Diego.
Treat our numbers as the shape of the cost, not the cost itself. We are explicit about this because most cost guides are not. If a figure is presented to you as precise, ask where it came from.
What our ranges are
- A realistic starting point before you call anyone
- A sense of what moves a price from the low end to the high end
- Enough context to tell a fair quote from an inflated one
- Free to read, with no phone number required
What they are not
- A quote, or a promise of what you will be charged
- Specific to your home, your market, or your condition
- A substitute for an in-person assessment on structural work
- Guaranteed current — pricing moves, and so do our figures
The evidentiary hierarchy we use
Not all published prices are evidence of the same quality. We rank sources by how close they sit to an actual transaction, and a figure’s verification status depends on which tiers support it.
| Tier 1 Transactional |
Prices a seller publishes and will honor: manufacturer and brand pricing, retailer listings including installed-price programs, and price lists published by working contractors. These are commitments, not estimates. A firm that posts a number it must then charge has an incentive to be accurate that a commentator does not. |
| Tier 2 Survey |
Aggregated data collected from practitioners under a stated methodology, such as the annual Cost vs. Value report. Weaker than a posted price for any individual job, stronger for establishing what a defined project costs nationally. |
| Tier 3 Derived |
Cost estimates published by sites that do not sell the work. Useful for triangulation and for spotting outliers. Never sufficient on its own. |
A figure is not marked verified on Tier 3 evidence alone. This matters more than it sounds, because Tier 3 is where most cost content on the internet lives, and a large share of it is published by lead-generation marketplaces whose pricing pages exist to move a reader into a quote form. That business model does not make their numbers wrong. It does mean the numbers were not produced to be accurate — they were produced to be plausible enough to keep someone reading until the form appears, and nobody is contractually bound by them afterward.
We read those sources. We use them to triangulate and to check whether we are an outlier. We do not count them as proof, and we do not cite them as authority.
We also weight convergence over volume. Three sites repeating the same figure is not three sources if the figure has one origin, and in this category it very often does. Two independent Tier 1 prices that agree are stronger evidence than a dozen Tier 3 pages that agree with each other.
Normalizing before comparing
The most common way a cost guide misleads is not by publishing a false number. It is by publishing a true number for a different thing. Before comparing any figure to a source, we normalize on three axes:
- Scope. Materials-only, installed, or turnkey. A materials price presented as an installed price understates by roughly the labor share of the job — frequently half of it.
- Unit. Per square foot, per linear foot, per room, per fixture, per visit, or a project total. Per-unit figures are only comparable after multiplying through by a stated typical quantity.
- Component versus system. The price of a deck surface is not the price of a deck; framing, railings, and stairs are part of the object a homeowner is buying.
Where a figure survives normalization and still disagrees with Tier 1 sources, it is an error. Where the disagreement disappears once the unit is corrected, it was a labeling failure — which we treat as an error too, because a reader cannot be expected to detect it.
Our first audit, and what it found
In August 2026 we ran the first formal audit of our own cost data, and we are publishing the result rather than describing the process and leaving you to assume it went well.
Design. We audited every headline cost figure in the library — the summary range that opens a cost article — 255 in total, after collapsing claims that appear on more than one page. Each was checked against sources retrieved at the time of the audit, under the hierarchy above, and classified with a fixed rubric. The audit drew on 528 sources, 418 of them Tier 1.
The twenty we owed you. An earlier version of this page reported 235 figures and withheld twenty. A research-tool limit had been reached late in the audit, and those twenty could not be checked against enough sources to judge fairly — reporting them as “unverifiable” would have credited a tooling failure as a finding, so we held them back and said we would re-run them. We have. All twenty were judged on the second pass, none came back unverifiable, and the results below now cover the full 255. The re-run went 8 sound, 9 too low, 3 too high — slightly worse than the set it joined, which moved our headline error rate up rather than down.
| Sound | Range overlaps Tier 1 or Tier 2 pricing, and any stated average falls sensibly inside it | 112 (44%) |
| Too wide | Not wrong, but spanning so broad a band that it carries little decision value | 18 (7%) |
| Too low | Materially below what published pricing supports | 106 (42%) |
| Too high | Materially above what published pricing supports | 17 (7%) |
| Unverifiable | No adequate published pricing found | 2 (<1%) |
Stated plainly: 123 of 255 figures — 48% — were materially wrong. Fewer than half did the job we publish them to do.
Precision of these estimates. At 255 figures the bands are reasonably tight: the 95% confidence interval on the materially-wrong rate runs from 42% to 54%, and on the sound rate from 38% to 50%. An earlier 25-figure pilot suggested 24% wrong; the full audit shows that was optimistic, partly because we tightened the rubric mid-audit to count ignored minimum charges as errors. Where our own earlier estimate and our later one disagree, the later and larger one stands. That has now happened twice: 24% became 47%, and 47% became 48% when the twenty withheld figures came back.
Error analysis
The errors are not random. 106 of 123 wrong figures are wrong in the same direction — too low — and they fail through the same mechanism: the bottom of the range is a materials-only, component-level, or best-case price presented as an installed, typical one. Examples include a composite decking figure that priced the deck surface rather than the deck, a fence figure below any installed price we could find for even the cheapest material, a gutter-cleaning figure below the minimum a company charges to send a truck, and an appliance service-call figure that turned out to be an hourly labor rate wearing a flat-fee label.
Sorted by frequency, the failures fall into five shapes:
| Average below the real floor | A stated average beneath the manufacturer’s own entry-level installed price |
| Minimum charge ignored | A per-unit price with no mention of the trade’s minimum call-out fee |
| Materials priced as installed | The cost of the thing, labelled as the cost of the job |
| Component priced as system | Part of the work priced as the whole of it |
| Cosmetic tier as full project | A refresh priced and labelled as a remodel |
A secondary pattern appeared in articles that passed: stated averages tended to sit at or below the floor of Tier 1 installed pricing. An average beneath the manufacturer’s own entry-level installed price is not an average.
What the statistics support. At the 25-figure pilot the directional bias was suggestive but not significant — a sign test returned p = 0.22, comfortably within what chance produces, and we said so. At 255 figures it is no longer in question: a sign test on 106 of 123 returns p = 3 × 10-17. The mechanism explains why. Unit substitution can only ever bias a figure downward, because the component left out always has a non-negative cost.
This distinction matters to us because it determines the remedy. Random error is fixed by checking figures one at a time. Systematic error is fixed by changing the rule that produced it — in this case, requiring that any range labeled “installed” have an installed price at its floor.
A second audit, on different content
A finding from one dataset can be a property of that dataset. The article audit told us our cost articles were biased low; it could not tell us whether the bias was in the articles or in how we produce cost data generally. So we ran the same method against a separate body of content that had never been examined: the cost tables on our service pages.
These matter more per figure than an article does. A service page’s cost table is the fallback source for every article beneath it that has no cost data of its own, so one wrong row propagates silently.
Design. The service audit is now complete. Of 112 service pages, 111 have been verified — every one except a single trade we could not source at all, discussed below. Each was checked against sources retrieved at the time, under the same hierarchy and the same rubric, drawing on 522 sources, 378 of them Tier 1 and 19 Tier 2.
| Cost summaries (111) | 27 sound · 61 too low · 3 too high · 20 too wide |
| Table rows (336) | 139 sound · 133 too low · 10 too high · 20 too wide · 34 unverified |
Stated plainly: only 27 of our 111 service cost summaries were sound. These are the headline figures at the top of every service page on this site, and three quarters of them were wrong or too vague to be useful.
61 of 64 wrong summaries were wrong low (p = 2 × 10-15), and 133 of 143 wrong rows (p = 7 × 10-29). Pooling both audits, 239 of 266 errors are too low, p = 7 × 10-44.
This page has now stated the directional claim four times and revised it twice, each time because a later wave produced counterexamples the earlier one had none of. We are leaving that history visible rather than presenting the current number as though it were always the number. The first 34 services produced no overstatements at all; the remaining 77 produced thirteen. A third audit, described below, produced nineteen more — the pattern of revision has not stopped.
The exceptions turned out to be the most useful thing in this audit, because they identify the real defect. Consider two of the three summaries that were too high:
- Our landscape design page priced landscape installation. Design is a professional fee — $70 to $150 an hour, or a flat $300 to $11,000 by scope, or 5 to 15% of construction cost. Two of the three rows in that table were the cost of building the landscape instead, a different transaction the reader may never make. The page was not overpriced so much as answering the wrong question.
- Our lawn fertilization page quoted $200 as a typical price. Every figure on it is a single-application price, but the summary reads as a project total. A real application is $47 to $90; a real annual program, which is what people actually buy, is $300 to $500. The published number sat in the gap between the two units and described neither.
Neither of those is an arithmetic error, and neither is really a story about direction. Both are scope and unit labels that do not say what the number means — and once you see the defect that way, the direction stops being mysterious. A range that quietly omits part of a job reads low. A range that quietly absorbs a larger job, or prices the wrong job entirely, reads high. The first is far more common than the second, which is why the bias is real and strong. But it is a consequence of the labelling defect, not the defect itself, and fixing labels is what we are actually doing.
The single most common form it took: a per-unit rate published where a project total belongs. Our interior painting page said “most projects run between $2 and $6,” because $2 to $6 is the per-square-foot rate. Our carpet cleaning page said $25 to $75, which is the price of one room. Our handyman page published its hourly rate as the cost of a job. Tile, wallpaper, exterior painting, lawn fertilization, window treatments, and snow removal all did versions of the same thing. Every one of those figures was true. Every one was in the wrong box, and the reader was told a house could be painted for six dollars.
The error classes ranked the same way as in the article audit, with scope errors dominating: materials priced as installed (14) and component priced as system (6) accounted for three quarters of them. Siding was the clearest case — vinyl published at $3 to $7 per square foot “installed” against $6 to $9 posted by a working contractor and $14.36 implied by a national survey. The page also priced wood siding below what its own fiber-cement row would have to be, a contradiction that needed no outside source to spot.
The one service we could not verify, and the mistake we made about it. Early in this work we piloted leak repair and ran roughly sixteen searches without finding a single Tier 1 or Tier 2 price: every contractor page that appeared to publish figures was restating the same handful of marketplace estimates, and several said so outright. That is false convergence in its purest form.
We then generalised from that one pilot and set aside fifty-nine services as unverifiable by nature — every trade that quotes after a diagnostic visit. That was wrong, and it cost us most of a category. When we later re-tested rather than assuming, appliance repair verified at Tier 1, then leak detection, then drain clearing, faucet work, water-line repair and foundation repair. In the end only leak repair itself stayed unverified.
The better test, which we now use, came out of the foundation work: a trade is verifiable when the unit is published, even if the count requires a site visit. Foundation piers are posted at $725 to $800 each; a job is a count of piers. Drain clearing is a flat rate per method. Faucet work sits in published price books. Leak repair fails because neither the unit nor the count is published anywhere — not because a plumber has to look first.
We are recording that error here because it is the kind a methodology page usually hides. Reasoning from category resemblance rather than evidence is exactly the failure this audit exists to catch, and we made it ourselves.
A third audit: every figure inside our cost guides
The first two audits checked headline figures — the summary range at the top of a page. That is the figure most likely to be read and quoted, but it is a small fraction of the numbers we publish. The rest sit inside the article: size tables, material comparisons, per-component costs, “for a 300 square foot job that works out to” sentences. Those had never been examined, and an earlier version of this page said so as a known limitation.
We have now closed that gap for one clearly defined body of content: every cost-type article on the site.
Design. This was a census, not a sample. Every dollar figure inside every article in the library — cost guides, how-to and comparison articles, and buying guides — was judged individually against the verified cost table above it: 13,624 figures across 1,574 article passes. Each figure had to be classified as sound, wrong, unverified, or not a price at all, and the classification had to account for every figure in the article. Nothing was skipped for being obviously fine.
| Sound | Agrees with the verified row that prices the thing the sentence describes | 8,524 |
| Wrong | Contradicts that row, or is labelled at a scope or unit it does not have | 2,923 |
| Unverified | No comparable row, and no Tier 1 or Tier 2 source found | 938 |
| Not a price | A threshold, deductible, rebate, permit fee or loan amount carrying a dollar sign | 1,239 |
Stated plainly: 2,923 of 12,385 price claims — 24% — were wrong. 734 articles needed at least one correction.
A bug in our own tooling, and what it cost. Partway through we found that the census had been skipping figures: it excluded any figure whose dollar string matched text we had written elsewhere in the same article, instead of excluding only the occurrences actually inside that text. 895 figures across 207 cost guides were never judged. We re-ran all 207 rather than estimate the damage — at the two-thirds mark the error rate was 1.8% and we could have stopped and published a confidence interval, but a census that extrapolates its last third is a sample wearing a census’s name. The true rate came in at 2.7%, above the point estimate we would have published. The skipped figures turned out to be overwhelmingly restatements of rows the verification had already passed, which is why the hole was nearly empty — but we did not know that until we counted.
The pilot underestimated this badly, and the reason is worth stating. A 40-figure pilot put the rate at 23–28%. The census came in at 31% overall and 39% in the harder half. The pilot was not wrong about the figures it sampled — it was wrong about the unit of measurement. Errors cluster inside articles. One mislabelled rate propagates into the size table, the comparison table, the derived total, the summary and the FAQ. Reading a whole article catches the cascade; sampling individual figures sees one link of it and cannot see that it is a chain. This is the fourth time an early estimate of ours has been revised by a larger one, and the fourth time it moved against us.
The finding that justifies the whole sequence
We ran the census in two waves, split by whether the service page above the article had been found sound or had needed correcting. The difference is the most useful number this audit produced:
| Under a corrected service page | Under a sound service page | |
| Articles | 399 | 136 |
| Figures judged | 4,746 | 2,162 |
| Wrong | 39% | 14% |
| Articles needing no change | 136 (34%) | 82 (60%) |
The mechanism, found later. Extending the census to general articles — how-to guides, comparisons, buying advice — showed what was actually propagating. Again and again, an article was still publishing its service page’s pre-correction cost table, verbatim, months after the service page itself had been fixed. Whole blocks of it: the summary range, the size ladder built on that range, the worked example, the FAQ. Sibling articles under one service would disagree with each other because some had been corrected and some had not. Of 574 wrong figures in general articles, 281 were this. The error was not written into each article independently; it was copied once and inherited everywhere, and correcting the source did not correct the copies.
An article sitting beneath an accurate cost table was nearly three times less likely to contain a wrong figure than one sitting beneath an inaccurate table. We had measured a version of this before — 37% versus 66% on whether an article needed any fix at all — and it has now reproduced on a different measurement, at a different scale, in the same direction.
One error class makes the mechanism visible. Unit mismatches — a per-unit rate written where a project total belongs — appeared 158 times beneath corrected service pages and twice beneath sound ones. That is not a coincidence of sampling. The article layer was inheriting the service layer’s defect rather than generating its own. Writers built size tables and worked examples on top of a headline number that was in the wrong box, and the error multiplied downward through everything derived from it.
The practical consequence is the sequencing rule we now work by: fix the anchor before auditing what hangs off it. Auditing articles beneath an unverified service table means checking them against a standard that is itself wrong, and the corrections you make will need making again.
Errors run in both directions
Everything above this point in our audit history pointed one way: figures too low, overwhelmingly, with a sign test that ruled out chance many times over. The census found the counterexamples, and we are giving them their own section rather than a footnote, because a bias this strong is exactly the kind of claim that becomes an assumption.
Two tree-service articles carried size ladders sitting a full tier above the verified table — a “typical job” at $2,600 against a verified average of $850. A further 17 figures in the second wave were stated averages above what published pricing supports. Same defect class as the low errors: a ladder or an average detached from its verified anchor. Opposite direction.
So the accurate statement of what we found is not “our figures were too low.” It is that our figures did not say what they were measuring, and the direction of the resulting error followed from which way the mislabel happened to go. Low is far more common, and the statistics on that are not in doubt. But direction is a symptom, and a reader who trusts a downward bias would have been misled by nineteen of our figures.
Buying guides fail in a way prices cannot show
Our buying guides — the “best X” articles — needed their own rules, because they recommend products rather than pricing jobs, and a product can fail a reader in a way a price cannot: it can stop existing.
So the check order was inverted. Before asking whether a price was right, we asked whether the thing was still sold. Across 115 buying guides that found 22 dead recommendations:
- Four makers had left the market entirely. One article recommended Panasonic solar panels under a “top efficiency” heading; Panasonic exited solar in April 2025. Another recommended Honda mowers; Honda no longer sells them in the US. A third named a lawn-care company that has traded under a different name for years. A fourth named a lighting company that no longer exists.
- Eleven product lines were discontinued and seven had been renamed or superseded.
Every price on those pages could have been perfect and the advice would still have been useless. That is the point: a cost audit cannot see this class of error at all, because there is no wrong number to find.
All twenty-two have now been replaced, and fixing them turned up three more the audit had missed — including a second Honda product, because Honda has also left the single-stage snow blower market. One flagged item was deliberately kept: a cordless snow blower that looked unavailable in late summer, where the manufacturer’s own catalogue and a dealer’s stock both said otherwise and an out-of-season listing gap had been mistaken for a discontinuation.
The blind spot in everything above
Every audit described on this page finds errors by looking for a dollar sign. That is what a cost audit is. It is also, we discovered, a serious limitation, and we would rather state it than let the volume of figures on this page imply a completeness we do not have.
While correcting solar figures, one of our reviewers noticed something sitting beside the numbers it was checking: the article assumed a 30% federal tax credit. So did sixteen others. That credit no longer exists. Both residential energy credits — the Residential Clean Energy Credit and the Energy Efficient Home Improvement Credit — were terminated for property placed in service after 31 December 2025. The IRS says so plainly on both credit pages, which we read directly rather than take second-hand.
It took us two passes, and the first one was not good enough. We found seventeen articles by searching for the phrase “30 percent” near “tax credit”, corrected them, and reported the job done. It was not. A second, wider search — for the claim rather than the wording — found nineteen more. They asserted the same dead credit as a bare dollar cap (“up to $600”), as a naked “Yes” in an FAQ answer, or as “federal incentives” with no number in the sentence at all. One solar article carried seven such claims and never once used a percentage.
For the days between the two passes, this site contradicted itself: four articles about the same heating and cooling equipment told readers opposite things about whether the credit existed. We are recording that because a reader could have landed on either one.
A second error was hiding underneath the first. Where the credit did apply, it was 30% but capped — $600 for a furnace, $600 for a central air conditioner, $2,000 for a heat pump. An article computing “30% of a $9,000 furnace, so $2,700 back” was wrong even in the years the credit existed. And roofing never qualified at all; that claim descends from a version of the credit that expired at the end of 2021.
None of this carries a dollar sign, so none of it was reachable by any audit on this page. An article can assert a terminated federal tax credit, in prose, on every page beneath a service, and a census of 13,624 dollar figures will pass it without comment. It surfaced because a reviewer flagged something outside their assignment rather than ignoring it.
We are stating this prominently because the same gap applies to claims we have not yet audited: permit and code requirements, contractor licensing and insurance, warranty terms, efficiency standards, and state rebate programmes. Those carry more risk per sentence than a cost range does, and they have not been through anything like this process. Two rebate figures in our HVAC articles are flagged internally as unverified for exactly this reason, and we have left them rather than quietly restate them as fact.
What our sources are
Across both audits we drew on 859 distinct sources spanning 524 domains. The composition matters more than the count:
| Tier 1 — transactional | Manufacturer and brand pricing, retailer installed-price programs, and price lists published by working contractors. The largest single contributor is a national home-improvement retailer’s installed-price catalogue. |
| Tier 2 — survey | Practitioner data under a stated methodology, principally the annual Cost vs. Value report. |
| Tier 3 — derived | Consumer cost-estimate sites. Read for triangulation, never counted as proof. |
We are not publishing the source list as links. A meaningful share of the Tier 3 corpus is lead-generation marketplaces, and linking out to them would send readers into exactly the quote funnels this page exists to give them an alternative to. Naming them as authorities would also overstate what they are: we read them, and we do not count them. If you want to check a specific figure, ask us for the sources behind it and we will send them.
A method that did not work
We first attempted to find these errors by script, on the theory that a mechanical pattern could be detected mechanically across the whole library. That attempt failed, and the failure is instructive enough to report.
The scanner flagged 344 articles. Hand-checking a sample of the flags found that nearly all were defects in the scanner rather than in the articles: two materials priced differently within one article were read as a self-contradiction, unrecognized units such as per board foot and per cubic yard were parsed as dollar totals, and articles opening with a table had a table row mistaken for their headline figure. The flag list was discarded in full.
A corrected pass produced a cleaner and more useful negative result. Across 242 headline figures, zero contained an internal contradiction — every figure agreed with every other figure in the same article. That internal coherence is why automated detection cannot work here: our figures are consistently wrong in the same direction rather than randomly inconsistent, and consistency is invisible to a consistency checker. Verification requires comparison against the outside world, one figure at a time.
We have since built several such scanners for different parts of this problem, and all of them behaved the same way. The best-designed compared each article’s figures against a service cost table already verified against real sources, so it was checking against known-good numbers rather than guessing. Across 1,172 articles put through per-article human-directed judgement, those scanners raised 5,382 flags. 2,448 were real. Forty-five percent, and that is the good version — the first scanner we built scored twenty-six.
The 2,934 rejections were not sloppy flagging. They were distinctions a person makes instantly and a pattern-matcher cannot make at all: a DIY kit price against a professional installed price, carpet quoted per square yard against a table in square feet, a pergola priced below a gazebo, a repair cost sitting under an installation service, a residential elevator measured against a stair-lift range, an outdoor condenser swap read as a whole air-conditioning system, the price of buying a washing machine against the price of fixing one.
We report this because the conclusion is load-bearing for everything else on this page. There is no cheap version of this work. A scanner can tell you where to look. It cannot tell you what you found, and a site that treats its flag list as a finding list will publish confident corrections that are wrong in a new direction.
The damage the corrections did
An article is not only its body text. It also has a short summary that appears on listing pages, a meta description that appears in search results, and a block of frequently-asked questions that search engines are entitled to quote directly. Those carry dollar figures too.
Two defects found during routine post-publication checking looked identical: in both cases we had corrected a price in the article body and left the old price standing somewhere else on the same page. We went looking for how often that had happened.
It had happened on 530 articles — a quarter of the library. And the pattern was not random. Articles our correction programme had edited carried a mismatch at 67%. Articles it had never touched carried one at 33%. Of 1,097 individual mismatches, 656 were an article whose body we had corrected and whose summary, meta description or FAQ block we had left byte-for-byte unchanged.
Some of that gap is explained by which articles we chose to audit — we went after the worst ones first, so they were more defective to begin with. That confound is real and we are not going to argue it away. But every rubric we wrote instructed the reviewer to check the summary, the meta description and the FAQs whenever they corrected a figure, and 656 untouched fields say that instruction was followed unevenly. This is a defect we introduced, not one we inherited.
The worst single instance: our tree-removal article’s summary was still telling readers the job costs “$500 to $5,000, with a typical job around $2,600” — the exact figure the first audit had removed from the body months earlier and replaced with $500 to $1,500 and a typical job of $850. Anyone who saw that article in a search result, or on a listing page, and did not click through, got the number we had already established was wrong.
Our deck-and-fence guide was worse in kind if not in degree. Its body warns readers in as many words to be suspicious of deck quotes in the $15 to $35 per square foot range, because that price covers decking and framing and leaves out the footings, railings and stairs that are roughly half the job. Its own FAQ block answered “How much does deck or fence installation cost?” with $15 to $35 per square foot. The article contradicted its own warning, in the field most likely to be quoted back to a searcher.
We corrected 199 articles across two passes — 132 summaries and meta descriptions, 67 FAQ blocks — conforming each field to the figure the body already carried, and checking that figure against the verified service table before doing so. No article body was changed by either pass.
Then we checked our own work on a live page, and found that the search had been suppressing the most common version of the defect.
The scanner skipped any field figure that sat entirely inside a body range, on the reasoning that a narrower claim is a legitimate refinement of a wider one. That reasoning is fine in the abstract and wrong here, because our corrections almost always widened ranges upward. A stale pre-correction figure nesting neatly inside its corrected replacement is not an edge case — it is the ordinary shape of a fix we half-applied.
The page that exposed it was the same tree-removal article. Its summary was fixed. Its FAQ still answered “How much does tree removal cost?” with “typically $200 to $2,000, averaging around $750” against a body reading $200 to $3,000 and $850. $200 to $2,000 sits inside $200 to $3,000, so nothing flagged it.
Re-running without that tolerance found 355 further mismatches across 226 articles, at a sampled accuracy of about 80% — the highest of anything in this programme. We corrected 143 more articles: 391 figures fixed, 132 rejected as legitimate, 155 of the fixes at the level of an article’s headline answer.
This is the same failure as an earlier one we publish below: a search whose result made the job look finished. We are reporting it because the alternative — quietly folding the extra 143 articles into the earlier total — would misrepresent how the work actually went. Three passes were needed, and the third only happened because we opened a page in a browser and read it.
Two corrections in that third pass were caught and rejected by an automated check before publication, because a reviewer had written a figure the article body did not contain: a solar headline given a $18,000 floor against a body saying $12,000, and a fence worked example implying $55 to $130 per post against a body saying $25 to $50. Both were re-done. We mention it because the check that caught them exists precisely because we assume our own reviewers make mistakes.
This produced a second negative result about scanners, and it is a more embarrassing one than the first. The FAQ flags were noisy, because a FAQ legitimately prices things the body never mentions — driveway markers, pool chemicals, the cost of renting a stump grinder — and no arithmetic can tell that apart from a contradiction. So we built a smarter filter to narrow them. The smarter filter was less accurate than the crude one it replaced: 15% of its flags were real against 33% for the scan it was meant to improve. It fired whenever a FAQ priced some component whose range happened to overlap the article’s headline range, which is common and almost always fine. We discarded it. What eventually worked was reading the FAQ question and keeping only answers to cost questions about the article’s own subject — 50% real, and the half that were real included every headline-level contradiction we knew about.
That is the second time in this programme that a more sophisticated filter performed worse than a blunt one, and both times we only found out because we sampled by hand before running it. A narrowing that is not measured is a guess.
We have left roughly 400 articles’ worth of lower-severity FAQ mismatches unexamined, and we would rather say so than imply the field is now clean. They are component-level — a sub-item priced slightly differently from the body — rather than headline contradictions, and finding them means reading every FAQ block in the library against its article. We also came out of this with 27 places, across 17 articles, where the article body disagrees with our own verified service table. Those are unresolved: some are article errors the census missed, and some are cases where the table is the stale side. We are not going to guess which.
Known limitations
- Non-numeric claims are largely unaudited, and that is the biggest gap here. See the section above. Tax treatment has now been checked for the seventeen articles that asserted a federal credit. Permit and code requirements, licensing and insurance claims, warranty terms, efficiency standards, and state rebate programmes have not.
- Every article type is now censused. That is not the same as everything being verified. 938 figures came back unverified — no comparable verified row, and no Tier 1 or Tier 2 source we could find. They are left in place and explicitly not counted as sound. They may well be right; they are unchecked, and we would rather label them than quietly promote them. We re-examined an earlier batch of them a second time against every service table rather than only the parent’s, which moved 43 into the sound column and turned up 2 further errors.
- Where a maker has left the market, the replacement is a judgement, not a lookup. All twenty-two dead recommendations are replaced, but choosing which brand now occupies a “best for X” slot is an editorial call and should be read as one. Where we could not source a reason for the new pick, we gave none.
- Our detection has failed more often than our judgement. Six times in this programme an automated search reported a count we then had to correct by reading: a stale-range scan that found 64 candidates and 14 real ones, a tax-credit scan that found 19 and meant 17, a cross-reference that suggested 92% of unverified figures were resolvable when the answer was 16%, a narrowing filter that scored worse than the crude scan it replaced, a tolerance rule that hid 355 mismatches by treating a stale narrow range as a legitimate refinement of the wide one that replaced it, and — the one that mattered — a tax-credit scan that missed nineteen articles because it required a particular phrase. Searches generate candidates. Only reading produces findings.
- Roughly 400 articles have unexamined FAQ figures. See the section above. They are component-level differences rather than headline contradictions, and we have not read them.
- 46 body figures across 30 articles disagree with our own verified service tables, found as a byproduct of the cross-field work and not yet resolved either way. Nine of them are on one page.
- One service — leak repair — remains unverified because no transactional pricing for it exists anywhere we could find. Its figures may be fine. They are unchecked, and its page says so.
- 34 individual table rows across the verified services came back unverified. They are not counted as sound.
- Our own claims about our error rate have been revised four times, each time by new data, and twice against us. Treat any single figure on this page as current rather than final.
- The rubric tightened mid-audit. Figures judged early were held to a slightly looser standard than figures judged late, which means the true error rate is more likely above our estimate than below it.
- Judgement is involved. On one near-identical pair of claims, two independent reviewers disagreed over whether a minimum service charge made a figure wrong. We added an explicit rule and re-ran, but some verdicts sit near a boundary.
- All figures are national. We do not currently publish regional adjustments, and regional variation is frequently larger than the error rates reported here.
- Published prices drift. A verification is a statement about a date, which is why we show the date.
What we changed
- Correction is underway against the sources that contradicted each figure. It is not finished, and we are not going to claim it is. To date we have rebuilt 84 service cost tables, corrected 1,044 articles whose headline or inherited figures were wrong, and applied 2,923 individual in-body figure corrections across 734 articles.
- We stopped copying cost tables between articles. The single largest error class in the general-article census was an article still publishing its service page’s pre-correction table, months after the service page was fixed. When a figure changes now, every article carrying a copy of it is changed in the same pass, and where sibling articles had drifted apart we converged them onto one wording rather than leaving several versions in circulation.
- We changed the order of the work. The census showed that articles beneath an inaccurate service table are nearly three times more likely to carry a wrong figure, so verification now runs anchor-first: a service cost table is verified before anything beneath it is audited against it. Auditing against an unverified standard means doing the work twice.
- When we correct a figure, we now correct it everywhere it appears on the page — the body, the summary a reader sees first, the search description, and the FAQ answers. We learned this the hard way: an early correction pass fixed article bodies only, which left the old wrong number sitting in the most prominent position on 49 pages.
- Any range labeled “installed” must now carry an installed price at its floor.
- Any per-unit price in a trade with a minimum call-out charge must state that minimum.
- A stated average must sit inside the range it summarises and above the cheapest real installed price we can find.
- Articles whose figures have been checked and corrected carry a dated verification line beside the numbers. Auditing an article is not enough to earn that line.
Corrections
If a number here looks wrong, tell us. We would rather fix a figure than defend it. Include the article and the number you are questioning, and if you have a quote in hand that contradicts us, that is the most useful thing you can send.
Corrections are not a formality. The most valuable evidence we receive comes from homeowners who have just been quoted a real price on a real house, because that is the one figure no published source can give us.
Common questions
Do you use AI to produce this content?
Yes. Our cost ranges and article content are produced with AI assistance, drawing on publicly available pricing information from across the home services industry. We say so plainly because we think you should know how the information you are reading was made. Verification, where it has happened, means a human-directed check of a figure against a named published source under the hierarchy described above.
Why publish your own error rate?
Because a cost guide that has never been audited and a cost guide that has been audited look identical until someone says which one they are. We would rather tell you that half the figures we checked were wrong, and what we are doing about it, than imply an accuracy we have not measured. Our own first estimate was 24%; the full audit came in at 47%; re-running the twenty figures we had withheld pushed it to 48%. On the figures buried inside articles, a pilot suggested 23% and the full census found 24% across 13,624 figures. Every time our number has moved, it moved against us, and each time we published the worse one because it is the true one. The most useful thing we have learned about our own estimates is that the optimistic ones were optimistic for a structural reason: errors cluster inside an article, so sampling individual figures always understates.
Which sources do you actually use?
859 distinct sources across 524 domains, so far. The largest contributors are a national home-improvement retailer’s installed-price catalogue, published price lists from working contractors, and the annual Cost vs. Value survey. We deliberately do not publish these as links: a large share of the cost content on the internet is produced by lead-generation marketplaces, and sending you into their quote funnels would defeat the purpose of this page. If you want to check a specific figure, ask and we will send you the sources behind it.
What does “cost data verified” mean on an article?
It means the cost figures on that page were checked against published pricing on the date shown, using the source hierarchy above, with at least one Tier 1 or Tier 2 source supporting the range. It does not mean the figures will match your quote — local pricing still varies.
How accurate are the numbers?
They are ballpark ranges, and we present them that way. They are useful for understanding roughly what a project costs and what drives the price. They are not accurate to your specific home, and no published range can be. Anyone quoting you an exact figure without seeing the job is guessing too — they are just not telling you.
How often is cost information updated?
We review cost information at least twice a year, and update figures sooner when a correction comes in or when an article is verified for the first time.
Why does your range differ from the quote I received?
Usually because of something specific to your job: local labor rates, access difficulty, the condition of what is being replaced, material grade, or how busy the contractor is. A quote outside our range is not automatically wrong — but it is worth asking the contractor to walk you through what is driving it.
Cost information is reviewed at least twice a year. Audits of record: 255 article headline figures (21–24 August 2026), 111 service cost tables, 336 rows (23–28 August 2026), a census of 13,624 in-body figures across 1,574 article passes (27 August – 1 September 2026), and a federal tax credit correction across 17 articles (30 August 2026), drawing on 859 distinct sources. This page last reviewed: August 2026.
