λ
signal
signal.lmbda.com
λ
signal • POST
The Internet Is Becoming More Synthetic — And We May Be Underestimating the Consequences
AI-generated text, images and automated accounts are rapidly expanding online. The deeper risks include eroding trust, synthetic consensus, distorted metrics and contaminated AI tr...
2026-08-30
Home / Technology & Society / Post
The Internet Is Becoming More Synthetic — And We May Be Underestimating the Consequences

The web was built on an implicit assumption: behind most pages, photographs, comments and videos was a person or organization attempting to communicate something. That assumption is beginning to fail. Generative AI can now produce convincing text, images, audio and video at a speed and marginal cost that human creators cannot match, while automated systems can publish and distribute those outputs continuously.

Fresh evidence suggests the shift is no longer marginal. In an August 2026 study, the Pew Research Center examined nearly half a million English-language webpages sampled from Common Crawl between 2021 and July 2026. In a random July 2026 snapshot, 10% of pages showed significant signs of AI authorship. More strikingly, when Pew looked only at pages published after ChatGPT’s November 2022 release, more than one-third showed signs of having been written or substantially edited by AI.

The exact percentages deserve caution. AI detectors are probabilistic rather than perfect, and Pew’s sample covers publicly accessible English-language pages represented in Common Crawl rather than the entire internet. But the direction is difficult to miss. The web is becoming a mixed environment in which human and machine production coexist at enormous scale, often without reliable labels telling users which is which.

The problem is not simply that AI content can be bad

Much of the discussion about “AI slop” focuses on obvious low-quality output: formulaic articles, impossible images, fabricated celebrity videos or endless social posts designed to harvest engagement. Those examples are visible because they fail aesthetically or factually. The deeper problem begins when synthetic content becomes good enough to blend into the background.

AI-generated material does not need to fool everyone to change the economics of publishing. It only needs to be cheap enough that producing thousands of pages, posts or videos becomes rational. Human creators traditionally faced a production constraint: writing, filming, illustrating and editing take time. Generative systems dramatically reduce that constraint, allowing publishers and automated accounts to test enormous quantities of content against recommendation and search algorithms.

Research into social platforms offers a glimpse of that dynamic. AI Forensics examined 354 automated or heavily automated TikTok accounts over one month in 2025 and identified more than 43,000 mostly AI-generated posts that collectively received 4.5 billion views. The organization reported that TikTok labelled less than 1.38% of the content in its dataset as AI-generated. The study is not a measurement of TikTok as a whole, but it demonstrates how inexpensive automated production can translate into enormous distribution.

When supply becomes effectively unlimited, attention becomes easier to manipulate

The internet’s major discovery systems were built to rank scarce human production. Search engines decide which pages deserve visibility; social networks decide which posts deserve distribution; recommendation systems decide which videos deserve another impression. Generative AI changes the supply side of that equation.

An operator no longer has to invest heavily in each individual piece of content. Thousands of variations can be generated, published and discarded while algorithms identify the handful that attract attention. This resembles automated experimentation in online advertising, except the experimental unit can now be the content itself.

The result could be an internet increasingly optimized by machines for machines. Generative systems create the material, ranking algorithms measure the response, automated operators adjust prompts or formats, and another generation of material appears. Human attention remains the scarce resource being competed for, but fewer humans may be involved in producing the things competing for it.

Search engines already recognize this risk. Google’s spam policies define “scaled content abuse” as producing large numbers of pages primarily to manipulate rankings rather than help users, explicitly including mass generation with AI when it adds little value. Google’s separate guidance says generative AI can be useful for research and structuring original material, while warning that generating many low-value pages can violate those policies. The distinction is important: synthetic production itself is not necessarily the problem; synthetic production at manipulative scale is.

Trust becomes expensive when fabrication becomes cheap

There is another consequence that is harder to quantify. As synthetic media becomes ubiquitous, authentic media can become less believable.

A realistic image of a disaster, an unusual animal encounter or a politician saying something inflammatory once carried some evidentiary weight simply because convincing fabrication required skill and resources. That friction is disappearing. The rational response is increased skepticism, but generalized skepticism has its own cost: genuine evidence can be dismissed as fake just as fake evidence can be accepted as real.

This is sometimes described as the “liar’s dividend” — the ability to exploit the existence of convincing synthetic media to cast doubt on authentic recordings. But the phenomenon extends beyond politics. Charities, journalists, scientists, witnesses and ordinary users increasingly need ways to establish provenance for material that previously could largely speak for itself.

That suggests authenticity may become an infrastructure problem rather than merely a media-literacy problem. Provenance standards, cryptographic signatures, platform labels and traceable publication histories can help establish where content came from. None is a complete solution: metadata can disappear, labels can be missed and sophisticated manipulation can occur after authentic capture. Yet an internet in which origin is routinely uncertain will require stronger mechanisms for establishing trust than an internet dominated by human production did.

Synthetic content may contaminate the systems that generate it

Perhaps the strangest feedback loop is that AI companies themselves depend on the web as a source of training data. If the web becomes increasingly synthetic, future models risk learning from material generated by previous models.

A 2024 study published in Nature investigated this recursively generated-data problem and described a phenomenon the researchers called “model collapse.” In their experiments, indiscriminately training successive generative models on model-produced data caused the learned distribution to degrade over generations, with less common parts of the original data distribution disappearing.

The result should not be interpreted as proof that using any synthetic data inevitably destroys an AI model. Carefully constructed synthetic datasets can be useful, and the Nature experiments explored particular recursive-training conditions. The important warning is about provenance and mixture. If large-scale web datasets become saturated with outputs from previous models and those outputs cannot reliably be distinguished from original human material, assembling high-quality future training corpora becomes harder.

The researchers argued that access to genuine human-generated data could therefore become increasingly valuable. In an ironic reversal, the mass production of synthetic information may make verified human information a scarcer resource.

The web’s measurements can become synthetic too

Content is not the only thing automation can manufacture. Views, comments, interactions, reviews and apparent consensus can all be influenced by automated systems. Generative AI makes this more sophisticated because accounts no longer need to repeat identical spam. They can produce endless variations in language, personality and imagery.

This complicates the signals people use to decide what matters. A post with thousands of enthusiastic comments may appear culturally significant even if many interactions are automated. A product may seem widely discussed because synthetic pages repeat descriptions derived from the same original source. A claim may appear independently corroborated because dozens of machine-generated articles reproduce it.

The danger is not simply false information. It is false independence. Repetition has historically acted as a rough credibility signal because multiple apparently separate sources implied multiple acts of observation or judgment. When content generation is automated, a hundred pages can ultimately trace back to one source, one prompt or even one error.

Human-created information may become a premium resource

The synthetic internet is unlikely to replace the human internet completely. A more plausible outcome is stratification. Cheap generated material will fill environments where volume is rewarded, while verified expertise, firsthand reporting, original research and recognizable human authorship become more valuable precisely because they are difficult to manufacture convincingly at scale.

This could change the economics of the web. Publishers may emphasize named authors, sourcing and reporting processes. Platforms may assign more weight to provenance and reputation. AI developers may pay more for curated human datasets rather than indiscriminately scraping public pages. Communities may retreat toward smaller spaces where identity and contribution history provide trust that open feeds cannot.

None of this means AI-generated content is inherently deceptive or worthless. Synthetic media can lower barriers to creativity, translation, education, software development and communication. Human creators can use generative tools while still contributing original judgment and expertise. The crucial distinction is not simply human versus machine; it is accountable information versus content produced without meaningful responsibility for whether it is true, useful or original.

Pew’s latest numbers provide an early measurement of a transformation already visible across the web. More than one-third of post-ChatGPT pages in its July 2026 sample showed signs of AI authorship. If production costs continue falling and autonomous systems become better at publishing and optimizing their own outputs, that share may keep growing.

The long-term consequence may be less dramatic than a “dead internet” populated entirely by bots, but more structurally important. The internet’s original scarcity was information. Its next scarcity may be confidence: confidence that a source is independent, that an audience is real, that an image records an event, that a recommendation reflects experience, and that the material training tomorrow’s machines still contains enough traces of the human world they are supposed to model.

Related
same category