The Label Travels
AI surveillance does not need to read your mind. It only needs to make suspicion cheap, portable, and hard to erase.
Conceptual illustration. The biometric match succeeds; an inherited record flag blocks clearance. Rahinah Ibrahim’s case involved an erroneous watchlist designation, not biometric failure.
In November 2004, an FBI agent filling out a watchlist form checked the wrong box beside Rahinah Ibrahim’s name. She was a Stanford doctoral student, not a terrorism suspect.
She learned what the box could do at San Francisco International Airport in January 2005. Officials later told her the mistake had been fixed. Yet related statuses kept changing in systems she could not see, and her visa was revoked weeks later.
On the first day of her trial in December 2013, Ibrahim’s American daughter, a listed witness, was denied boarding in Kuala Lumpur after Homeland Security sent the airline a “possible no-board request.” The district court called that episode a mistake and said it was quickly remedied. The public record does not establish that Ibrahim’s original erroneous designation caused her daughter’s denial.
After more than nine years, the government conceded at trial that Ibrahim was not a threat. A federal court ordered the mistaken designation cleansed across its records. The error took a moment to make. It took a lawsuit and a trial to unmake.
Her case predates today’s AI systems. Its administrative logic does not.
The familiar nightmare is a state that hears everything and understands everything. The more durable danger is duller: a low-quality label attaches to a person, then moves from database to database long after anyone has stopped checking the evidence beneath it.
The danger is not comprehension. It is portability.
The first pass became cheap
Agencies have filtered intercepted material by machine for decades: keyword dictionaries, known selectors, contact chains. In disclosed bulk systems, collection and automated reduction often outran human review. Analyst attention became a central constraint, although interception access, encryption, language, storage, identity resolution, and law remained constraints too.
The current wave of AI lowers the marginal cost of a first pass. It can transcribe, translate, link, rank, and revisit material without asking a person to begin with every item.
A 2023–24 Florida corrections contract shows the new price of that first pass. The state bought automated transcription and search of recorded prison calls at five cents a minute, with an annual ceiling of fifty million minutes. The public record does not say that every minute under the ceiling was processed. But fifty million minutes, heard without interruption by one person, would take about ninety-five years.
Where the ambition reaches a population, the evidence becomes thinner. A leaked NSA presentation dated June 2012 described preliminary behavioral classifiers run over more than 55 million Pakistani phone selectors. The initial positive set contained only seven known courier selector pairs; later work added weak labels. Selectors are not people, and the slides do not establish a validated production system.
The next step is not simply a larger shortlist. A model can summarize a large corpus placed within its reach and draft a fluent memorandum around the result. In May 2026, the Pentagon said it had entered agreements with eight firms to deploy advanced AI capabilities on its IL6 and IL7 classified networks. Anthropic was absent after a documented disputeover proposed limits on mass domestic surveillance and fully autonomous weapons.
That procurement establishes classified access, not a public record of population profiling. It does point to a quieter risk. A flag accompanied by polished prose may be easier to forward and harder to question. The memorandum can look like verification when it is only the ranking, rendered in sentences.
The systems still cannot verify belief. They sort observable proxies and leave institutions to supply the inference.
When Human Rights Watch reverse-engineered the Android app associated with Xinjiang’s Integrated Joint Operations Platform, it found fields for electricity use, movement, phone identifiers, foreign contacts, and flagged network tools. HRW could not log into the central server, so the audit revealed an apparent workflow rather than every live rule.
A separate Aksu detainee list, analyzed by HRW in 2020, recorded tools including Zapya and VPN software among reasons associated with people sent to political education. That second source connects IJOP flags to actual detention decisions, although it still does not show that an algorithm made those decisions alone.
The system records conduct and identifiers; officials supply the ideological meaning. A foreign call, an app, a movement, or an association can therefore acquire a significance the underlying observation did not contain.
The arithmetic of error
The numbers impose a harder limit. Screen a whole population for a rare condition and even a seemingly strong classifier can produce a queue dominated by false alarms.
Take a modeled illustration: a trait present in one person in 100,000, a population of 200 million, a system that catches 99 percent of true cases, and a 1 percent false-positive rate among everyone else. A single pass produces 1,980 true alerts and 1,999,980 false ones. Of 2,001,960 alerts, only 0.099 percent—about one in 1,011—is true. False alerts are not a marginal error in this model. They are 99.9 percent of the output sent for review.
Modeled illustration, not measured performance of a named system. Assumptions: one independent screen per person, prevalence of 1 in 100,000, sensitivity of 99 percent, and a false-positive rate of 1 percent. Alerts may not represent unique people or account for repeated screening, subgroup variation, or human review. Conceptual background: National Research Council, 2008.
Compute does not dissolve the base rate. Nor does bulk collection guarantee useful results. After a classified review of one specific NSA program—the bulk domestic telephone-metadata program—the Privacy and Civil Liberties Oversight Board’s majority found no case in which it made a concrete difference to the outcome of a counterterrorism investigation. It found limited value in corroborating information obtained elsewhere.
Automation can still be useful. It can narrow a corpus, route a language, or surface a known pattern. What it cannot do is repeal the mathematics of rare events. At population scale, the operational gain is a shorter queue. The accuracy of what follows remains a separate question.
When a flag becomes a fact
The box beside Ibrahim’s name became powerful because it traveled.
Once a person is categorized, the category becomes a small administrative fact that can be copied from one institution to the next at almost no cost. Handoffs create an asymmetry. Honoring an inherited flag may be easier than revalidating or removing it, especially when the recipient has no authority over the upstream record.
Ibrahim documents that problem in one watchlisting system. It does not prove identical incentives at every border, bank, police department, or employer. But the cost of error often lands on the person least able to see the file, while the receiving institution sees only the status it was given. A provisional signal can harden as it moves.
China has built an increasingly standardized lattice of court and sectoral blacklists. It is real, but it is not the mythical single AI-generated “social-credit score” assigned to every citizen. The court-defaulter system, for example, attaches specified travel and consumption restrictions to defined enforcement categories. Some local points systems exist, but they are not one national score.
The cleanest measured example needs no algorithm. Venezuela’s 2004 Maisanta database identified people who had signed a petition seeking a recall referendum against Hugo Chávez. A later matched study estimated that identified signers suffered, on average, about 5 percent lower earnings and a 1.3 percentage-point fall in employment after the list spread. A category crossed institutional boundaries, and the economic penalty followed.
The Western cases are less integrated. No public evidence shows one engine linking policing, finance, employment, and education into a universal score. What exists is a collection of narrower systems with their own laws and data.
In Britain, a Prevent referral is a safeguarding process, not a conviction. Counter Terrorism Policing generally keeps the referral record until a scheduled review six years after closure, including cases marked “No Further Action” that never reach the Channel program. Records may remain longer when a policing purpose is documented.
In the United States, the State Department expanded online-presence review for specified student and exchange-visitor visa applicants in 2025 and required covered applicants to make social accounts public. Axios reported, citing senior State Department officials, that a separate “Catch and Revoke” effort used AI-assisted reviews of existing student-visa holders. The public record does not disclose the tools, error rates, or how many later revocations resulted from those reviews. Effective March 30, 2026, the State Department extended online-presence review to fourteen further nonimmigrant classifications.
A DHS Inspector General audit confirmed that Customs and Border Protection, Immigration and Customs Enforcement, and the Secret Service procured and used advertising-ID-derived mobile-device location data in fiscal years 2019 and 2020. ODNI separately acknowledges intelligence-community access to, collection of, and processing of commercially available information more broadly.
The public record does not establish that those purchases feed every watchlisting or screening system described here. The architecture is fragmented. The common feature is narrower: a classification made in one context can become useful to another institution before the person concerned has any practical way to inspect it.
The best-known mislabelings are old. Ibrahim’s box was checked in 2004; the Maisanta database circulated that year. That is not evidence that newer systems are cleaner. Comparable errors may remain invisible because affected people are often not notified, and the public record often does not disclose a denominator. A person usually discovers a hidden classification only when it produces a visible consequence.
Seven years for a car payment
One design principle follows from these cases. If a label is allowed to travel, its source, confidence, expiry date, and route of appeal should travel with it.
The United States built a version of that principle for money. Under the Fair Credit Reporting Act, adverse-action notice identifies the reporting agency; the agency must disclose the consumer’s file; a dispute generally triggers reinvestigation within thirty days; and most adverse information expires after seven years. The regime is imperfect and litigated, but its architecture is visible: provenance, access, correction, and time limits attach to a portable record.
National-security watchlisting has redress routes, including DHS TRIP. It has no comparable general statutory package of routine provenance, confidence, fixed expiry, file access, and correction. Ibrahim’s mistaken designation continued to affect related records until a court ordered correction more than nine years after the original nomination.
Comparison of statutory architectures. Individual programs and remedies differ. The credit-reporting lane summarizes general FCRA protections; the security lane does not imply that every program lacks notice or redress.
Security files are not credit reports, and sources cannot always be exposed. Yet an expiry date names no informant; a confidence grade need not reveal a method; and a route of appeal discloses only a door. Secrecy can protect a source while still allowing the label itself to be checked.
A missed payment generally leaves a credit report after seven years. Ibrahim needed more than nine years and a federal trial to correct a designation she was not permitted to inspect.
The limits
The techniques increasingly resemble one another across political systems: filtering, graph analysis, entity resolution, data fusion, ranked triage. What differs is everything around them—who may authorize collection, for what purpose, under what oversight, and whether a person can ever inspect and contest the file.
That legal layer constrains power, but it is neither simple nor permanent. Title VII of FISA, including Section 702, automatically repealed on June 12, 2026. A transition clause nevertheless preserved existing orders, authorizations, and provider directives until their stated expiration dates, publicly described as roughly March 2027. No new annual certification can issue without legislation.
Transparency is also partial. Current federal guidance exempts intelligence-community elements and the Department of Defense from the government-wide public AI use-case inventory. That means the public cannot calculate how much AI adds over the older machinery of lists, selectors, and database joins.
Courts sometimes narrow collection. On August 5, 2026, U.S. District Judge Carlton Reeves upheld refusals to issue four original and three revised tower-dump applications and held the technique per se unconstitutional as a general warrant. It is one district-court order in a divided field, not a national rule.
Encryption remains a physical counterweight. Properly implemented end-to-end encryption can keep message content unreadable in transit, although metadata, compromised endpoints, backups, and commercial location records may remain available. Surveillance does not become omniscient merely because classification becomes cheaper.
What outlives the file
The damage was never limited to accurate control.
East Germany’s Stasi ran no prediction engine. Yet a peer-reviewed study linked higher average informer density in East German counties during 1980–88 to lower post-reunification trust, civic participation, and income. The design cannot isolate whether fear, repression, or the quiet shredding of ordinary relationships did the work.
It is a historical analogue, not proof about AI. It carries a warning nonetheless. A system does not have to understand a person, or even be right about that person, to change how people speak, whom they call, and what they are willing to sign. It only has to make a label cheap to produce and costly to shed.
Rahinah Ibrahim’s case leaves a narrower question: how long should an unverified designation be allowed to travel before someone must check it? American credit law gives financial labels correction and expiry. Suspicion has no comparable general regime.
Her designation moved through government systems for more than nine years before a court forced correction.
What would have changed if its uncertainty had been required to travel with it?
Selected sources
• Ibrahim v. Department of Homeland Security, Ninth Circuit opinion, January 2, 2019
• Florida Department of Corrections, VERUS contract C3086
• NSA SKYNET presentation, June 5, 2012, published by EFF
• Human Rights Watch, reverse engineering the IJOP app, May 1, 2019
• Human Rights Watch, Aksu IJOP detainee list, December 9, 2020
• National Research Council, Protecting Individual Privacy in the Struggle Against Terrorists, 2008
• PCLOB, report on the Section 215 telephone-records program, January 23, 2014
• UK Home Office, Lessons for Prevent, updated November 3, 2025
• U.S. State Department, expanded online-presence review, effective March 30, 2026
• DHS Office of Inspector General, commercial telemetry data audit, September 28, 2023
• ODNI, Commercially Available Information Framework, May 2024
• Federal Trade Commission, Fair Credit Reporting Act
• Congressional Research Service, status of FISA Title VII after June 12, 2026
•Lichter, Löffler, and Siegloch, “The Long-Term Costs of Government Surveillance,” JEEA, 2021





