In short
Lighthouse scores vary because the test measures timing on whatever machine ran it, under simulated throttling, against a page whose third-party scripts, ads and server response differ every load. Variance of five to ten points between runs is normal; larger swings usually mean a busy machine, browser extensions, or a heavy third-party. To get a usable number, run five times in an incognito window with extensions off and take the median, or use PageSpeed Insights, which runs on consistent hardware. To know whether users are affected, use the field data instead.
What Lighthouse is measuring
Lighthouse loads the page in a controlled browser, simulates a slower device and network, and records timings: when the first content painted, when the largest element rendered, how long the main thread was blocked, how much the layout moved. Those timings are converted to a score. Every one of them is a measurement of a process that involves your machine, your network, the page's server, and every third party the page loads, and none of those behaves identically twice.
So the score is a sample, and samples vary. The question is how much variance is noise and how much is signal.
The seven sources of variance
1. The machine running the test. Lighthouse's mobile simulation throttles CPU by a factor relative to the host machine. A laptop that is also running a video call, a build, or forty browser tabs produces slower timings and a lower score. This is the largest source of local variance and the least noticed.
2. Browser extensions. Ad blockers, password managers, and developer extensions inject scripts and alter page loading. Chrome's Lighthouse panel warns about this; many people ignore the warning.
3. Network conditions. Simulated throttling is applied on top of the real connection. A congested wifi or a VPN adds real latency that the simulation does not remove.
4. Server response. Your server may answer in 80 milliseconds when warm and 800 when cold, when the cache has expired, or when a background job is running. Time to First Byte then shifts everything downstream.
5. Third-party scripts. Analytics, tag managers, chat widgets, consent tools, embedded videos and ads all load from other people's servers at other people's speeds. An ad slot that serves a heavy creative on one run and a light one on the next can swing the score by fifteen points on its own.
6. A/B tests and personalisation. If the page serves variants, Lighthouse may be measuring a different page each time.
7. Score thresholds. The score is derived from metric values using curves, and near a threshold a small change in one metric moves the composite noticeably. A page hovering at 89/90 will flip colours between runs for no real reason.
Google's own guidance acknowledges the variance and recommends multiple runs; the Lighthouse team has documented that the performance score can vary by several points on identical pages.
Getting a number you can trust
- Use an incognito window with no extensions, and close everything else on the machine.
- Run five times and take the median score, or better, the median of each underlying metric. Lighthouse CI and WebPageTest can automate this.
- Use PageSpeed Insights for a consistent environment. It runs Lighthouse on Google's infrastructure, which removes your machine and your network from the equation. Its results still vary, but less, and for the reasons in the page rather than the tester.
- Compare metrics, not the score. Largest Contentful Paint in milliseconds is more stable and more meaningful than the composite. When judging a change, look at whether LCP, Total Blocking Time and Cumulative Layout Shift moved, not whether the number went from 84 to 87.
- Test the same URL, logged out, with the same cache state. A warm cache and a cold cache are different tests.
- Test at the same time of day if the server's load varies.
Lab versus field
All of the above is lab data: a single synthetic load under fixed conditions. It is useful for diagnosis because it explains itself and it is reproducible enough to compare before and after a change.
The number Google uses for ranking is field data: Core Web Vitals measured from real Chrome users over 28 days, reported at the 75th percentile. Field data does not have run-to-run variance because it is an aggregate of thousands of loads, and it reflects real devices in real places. It is available in Search Console and at the top of PageSpeed Insights for pages with enough traffic. When lab and field disagree, the field is right about what users experience, and the lab is right about what to fix.
Setting up real-user monitoring on your own site gives you field data for every page, not just the popular ones, and removes the temptation to argue with a single Lighthouse run.
When a change is real
A difference between two single runs is not evidence of anything. A difference between the medians of five runs each, on the same machine under the same conditions, of more than a few points, or of more than a couple of hundred milliseconds on LCP, is probably real. A change in the field data over the following weeks is definitely real. Judge optimisation work by the last of these, and use the first only to decide what to try.
If your score swings wildly and you cannot see why, the usual culprit is a third-party script; load the page with each one blocked in turn and watch the variance disappear. If it does not, book a call and we will find the source together.
Common questions
Why does my Lighthouse score change every time I run it?
Because it measures a timed simulation on your machine, and the machine's load, browser extensions, network conditions, the server's response time and the page's third-party scripts differ on every load. Variance of five to ten points is normal. Run several times in a clean incognito window and use the median.
Why is my Lighthouse score different from PageSpeed Insights?
PageSpeed Insights runs Lighthouse on Google's servers with consistent hardware and network, whereas your local run reflects your own machine and connection. PageSpeed Insights also shows field data from real users, which local Lighthouse cannot. Treat the PageSpeed Insights lab score as the more comparable of the two.
How many times should I run Lighthouse?
At least three and preferably five, taking the median of the score and of each metric. Tools such as Lighthouse CI and WebPageTest automate repeated runs. A single run is a sample and should not be used to judge whether a change helped.
Which Lighthouse score matters, mobile or desktop?
Mobile. Google indexes and ranks the mobile version of a site, most visitors are on phones, and the mobile simulation's throttled CPU exposes JavaScript problems that a desktop run hides. Desktop scores are useful for diagnosing desktop-specific layouts but should not be the headline number.
Does Google use the Lighthouse score for ranking?
No. Google ranks on Core Web Vitals measured from real Chrome users over 28 days, reported in Search Console and in the field data section of PageSpeed Insights. Lighthouse is a diagnostic lab tool; improving what it reports usually improves the field numbers, but the score itself is never seen by the ranking systems.
