Data completeness is whether the records you have represent everything you should have. It is the hardest data quality dimension to measure, because it requires knowing about things you do not have.
This is the part worth internalising. Missing data almost always makes security metrics look better, not worse.
Assets the scanner cannot reach do not appear as unpatched. Endpoints where the agent failed to install do not appear as uncovered. A business unit that has not onboarded contributes no findings. A source that stops reporting stops contributing failures.
The bias is systematic and in one direction. A metric programme with an unmeasured completeness problem is a metric programme reporting optimistically, consistently, without anyone acting in bad faith.
Against an independent population. Compare the tool’s view against an inventory built from a different method. The disagreement is the estimate. See source of truth.
By expected volume. Each source has an expected record count per period. Deviation from it is a completeness signal before it is anything else.
By coverage of the known universe. Share of business units, regions, environments and asset classes actually contributing data. A metric drawn from four of seven regions is a regional metric being reported as an enterprise one. See coverage gap.
A coverage figure of 96 percent computed over 60 percent of the estate is a different claim from the same figure computed over 98 percent.
Publishing the completeness of the underlying population next to the value costs one line and prevents the most common misreading of a security metric.
Sometimes completeness cannot be established, because there is no independent view of the population. Unmanaged cloud accounts and shadow IT are the standard examples.
The correct response is to state the limitation rather than to report the metric as though the population were known. A number with a stated boundary is usable. A number implying a boundary it does not have is not.
From the blog