You could just ask.
So why didn't it say?
Every page on this site could be replaced by a question typed into a chat window, and the answer would be fast, fluent, and mostly correct. That objection deserves a real answer, not a wave of the hand. It survives right up until you ask what the answer is made of, what it leaves behind, and where it was computed. The last one turns out to matter most.
Grant the objection in full
Ask a competent model what Flock Safety sells, or what Palantir's role in federal data integration is, and you will get a serviceable answer in about four seconds. It will be better organised than most journalism on the subject. It will be free. It will not condescend to you, and it will let you ask the follow-up question you were too embarrassed to ask a person.
That is a real improvement and it should be said plainly, because the argument only works once it is conceded. This site is not competing with that and would lose if it tried. Nothing below is an argument that the machine is stupid or that you should not use it.
The argument is narrower. A generated answer is a very good instrument for a question you already have. Almost none of the material on this site is something anybody was going to ask about.
The hard part was never the explaining
Consider what it would take to arrive at the page on your car by asking. You would have to already suspect that a driver-facing camera requirement exists, which almost nobody does, because it did not arrive as a product announcement. You would have to know it was worth asking about before you had any reason to think so.
Same with age verification. The question most people would think to ask is whether a given law is good or bad. The thing actually worth knowing is what your phone permanently holds after you prove your age once, which is a question you can only formulate if somebody has already told you that is where it ends up.
A model answers questions. It does not hand you an inventory of the things you have no idea are happening. That inventory is the entire product here, and the explanations attached to each item are the cheap part, the part that could indeed be regenerated by anyone in seconds.
It will produce a plausible list, weighted towards whatever was written about most, which means weighted towards the stories that already got attention. The items that are underreported are underreported in the training data too. A generated survey of a field reproduces that field's blind spots with total confidence and no visible seam.
This is the failure mode that matters and it is invisible by construction. You cannot see what a fluent answer left out.
An answer with no address
Every substantive claim on this site is supposed to arrive with three things attached: the document it came from, the date it was checked, and an admission when the document does not actually say what would be convenient. That is not a stylistic preference. It is the only mechanism by which a reader can find out that this site is wrong.
A generated answer typically has none of the three. It is not lying about its sources, it simply does not have any in the sense that matters. There is no filing to open, no docket to pull, no date on which the underlying fact was true. You are given a conclusion and no way to audit it, which is a fine arrangement when the stakes are a recipe and a poor one when the claim is that a specific company holds a specific power over you.
The tables on this site say when they were run and tell you to re-run them rather than trusting them indefinitely. That instruction separates a number you can check from one you have to believe. It is not humility for its own sake.
The point is not that this site is reliable and a model is not. Anything written by a person can be wrong too, and some of what is here will turn out to be. The difference is that when this site is wrong, the wrongness has an address: a document you can open, a date you can check, and somebody you can tell.
Ask twice and nobody can tell you what changed
Ask the same question in six months and you may get a materially different answer. Nothing announces this. There is no changelog, no diff, no record that the earlier answer was ever given. If a model's handling of a politically awkward subject shifts, whether through deliberate adjustment or as a side effect of retraining, that shift is not observable to the person asking.
Every word on this site is in a public repository with a commit history. Anyone can see what a page said last year, when it changed, and what the change was. That is a modest property and an unglamorous one. It is also the only reason a reader has to treat a correction here as meaningful, because a site that can silently rewrite its own past is not making claims, it is producing weather.
Ephemerality is not a bug in a chat interface. It is the right design for a conversation and the wrong substrate for a record, and most of what is worth knowing about surveillance is a record.
The question is also an answer about you
Typing a specific worry into a hosted service produces a record of that worry, held by somebody else, on terms set out in a policy you did not negotiate. This is true of search engines and has been for twenty years, so it is not novel. What is different is the specificity. A search query is a fragment. A conversation is a narrative, in your own words, about what you are afraid of and why, often including the details you were reluctant to say out loud.
Be careful about how far to push this. Retention terms vary by provider and by plan, some offer configurations that do not retain, and there is a real difference between a company holding data and anybody looking at it. Overstating this is exactly the kind of unsourced scare claim that makes the careful version easier to dismiss.
The accurate version is narrow and sufficient: researching surveillance through a hosted service is not a neutral act, it is a disclosure, and you are better off knowing that. A static page you read produces no such record, because there is nothing here to produce it with.
Where the answer is computed
This is where the media argument becomes a surveillance one.
The answer you get did not come from your device. It was computed in a building, and that building is the same class of building, frequently operated by the same handful of companies, that makes surveillance at population scale economically possible. That is a statement about capital expenditure, not a metaphor about how everything is connected.
Mass surveillance has never been limited by cameras. Cameras are cheap and have been for a decade. What limited it was the cost of keeping everything and searching it later. A plate reader that stores thirty days of data is a tool for investigating a specific crime. The same reader feeding storage that keeps several years, indexed so that any officer can ask where a given plate has been since 2023, is a different instrument entirely, and the only thing separating the two is how much compute and storage cost.
That cost has been collapsing, and it has been collapsing because an enormous amount of money is being spent building exactly the capacity that makes it collapse. The buildout is not being justified by surveillance. It is being justified by demand for the thing that answers your question in four seconds. The capacity does not care which it is used for.
The scale is now large enough to show up in national energy statistics. In a report released on 20 December 2024, the Department of Energy found that data centres consumed about 4.4% of total U.S. electricity in 2023, and projected they would consume approximately 6.7% to 12% by 2028.
A doubling of an entire sector's share of national electricity inside five years is construction, not a software trend, and construction is the one part of this that cannot be done quietly.
The building files paperwork the software never does
This is the practical reason to watch the physical layer, and it is close to the opposite of what most people expect.
A model's training data is a trade secret. Its retention policy is a document the company writes and can change. Its content decisions are unauditable from outside. But a building that draws tens of megawatts has to ask somebody for the power, and that request generates a public record. Siting requires zoning. Zoning requires hearings. Water withdrawal requires permits. Tax abatements require a vote by named people at a published meeting. Federal facilities generate environmental assessments.
The physical substrate of this industry is dramatically more legible than the software running on it, and it is legible through exactly the local, boring, procedural channels described on the contribute page. Even the classified end of it leaves traces: the Army Corps of Engineers published a draft environmental assessment for a campus expansion at the National Security Agency's Utah Data Center in February 2025, because building the expansion requires saying so in public.
The same visibility applies to who is buying. Federal contract obligations are public and queryable. The figures below come from a query run against the U.S. Treasury's USAspending API (opens in a new tab) on 18 August 2026, covering prime contract awards from 1 October 2022 to that date, matched on recipient name.
| Recipient | Obligated | What to read into it |
|---|---|---|
| Microsoft Corporation | $1.94 billion | Spans cloud, licensing, and software. Not separable into infrastructure alone from this query. |
| Amazon Web Services, Inc. | $1.05 billion | A further $54.7 million sits under a variant spelling of the same name. |
| Oracle America, Inc. | $217 million | Combined across two spellings of the same entity. |
| Google Public Sector LLC | $47.9 million | Google LLC adds a further $15.9 million. |
Read this table narrowly, because it is narrow. These are obligations recorded against prime contract awards, not revenue received, and they exclude anything bought through a reseller, which is a substantial share of federal software purchasing. Name matching is crude: the same query returns $4.05 billion to Oracle Health Government Services, which is electronic health record work rather than cloud infrastructure and is excluded above, and it also returns an unrelated lift company that happens to have Oracle in its name. That is a good illustration of why a figure from a search result is worthless. Re-run the query rather than trusting these numbers indefinitely.
The four companies above are also the four awarded the Defense Department's Joint Warfighting Cloud Capability in December 2022, a single vehicle through which the department buys commercial cloud across classification levels. The overlap is the point. The infrastructure serving your question and the infrastructure serving a defence department's data are, at the level of concrete and transformers, the same industry.
What actually changes things
- Data centre siting is decided locally, by people you can reach. County commissions, planning boards, and utility regulators approve these facilities. The hearings are public, scheduled in advance, and usually attended by nobody. This is the most reachable decision point in the entire industry and it is currently uncontested.
- Ask what the abatement bought. Facilities are routinely granted tax relief on a jobs argument. The permanent staffing of a data centre is small by design. Whether that trade was worth it is a legitimate question with a checkable answer, and the agreement is generally a public document.
- Ask about the power and the water in the same breath. Both are finite locally, both are allocated by regulated entities, and both create a public record. A utility interconnection request is one of the few reliable early signals that anything is being built at all.
- Keep the retention question separate from the capability question. Arguing about whether compute is good is unwinnable. Arguing about how long a specific system keeps a specific record, and who may query it, is concrete and has repeatedly succeeded.
None of this requires believing that anybody involved is acting in bad faith, and the argument is weaker if it does. The companies building this capacity are building it to sell a service people want. The agencies buying it are buying it for reasons they will state on request. The model that answers your question in four seconds is genuinely useful and will keep getting more so. What follows anyway, without anybody deciding it, is that the cost of keeping everything about everyone and searching it years later falls every quarter, and that the institutions best positioned to use that capability are the ones already buying the capacity. That is not a prediction, and not a conspiracy either. It is a line item in a public database that anybody can query. The reason this site exists is that almost nobody does, and a summary generated on demand will never tell you to go and look.