Close Menu
    Facebook X (Twitter) Instagram
    TRENDING :
    • What happened when Meta tried to collect employee data
    • Georgia High School Pulls Off Stunning Last-Second Win
    • Don’t trust AI companies with your content. Verify them instead
    • Jacory Croskey-Merritt Fantasy Week 1: Volume Play at Flex
    • The iPhone Duo isn’t just a folding phone
    • Patrick Mahomes Fantasy Week 1: Start Your First-Round Pick
    • I bought a Birkin as an investment. It taught me a hard lesson about human judgment
    • Jalen Hurts: Start Week 1, Then Trade
    Populist Bulletin
    • Home
    • US Politics
    • World Politics
    • Economy
    • Business
    • Headline News
    Populist Bulletin
    Home»Business»Don’t trust AI companies with your content. Verify them instead
    Business 7 Mins Read

    Don’t trust AI companies with your content. Verify them instead

    Business 7 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Telegram Email Copy Link
    Follow Us
    Google News Flipboard
    Share
    Facebook Twitter LinkedIn Pinterest Email

    If it’s a day ending in Y, you can count on a story that further erodes the public’s trust in Big Tech. This week it was Sony Music and Warner suing Anthropic over copyright, accusing the AI company of illicitly pirating music and song lyrics from its catalog to train its AI models. With the lawsuit, Anthropic is now officially the target of all three major music publishers, since Universal brought a similar action in January.

    As copyright lawsuits go, the allegations are pretty juicy. They paint a picture of Anthropic brazenly pirating music catalogs through torrenting and copying huge troves of lyrics from third-party websites wholesale. The filing claims Anthropic cofounder Benjamin Mann personally conducted or directed the torrenting and discussed it openly in Slack channels. Anthropic, which agreed a year ago to pay $1.5 billion in a settlement over pirated books, responded curtly, telling Axios “we intend to defend ourselves robustly in court.”

    To any media executive, content creator, or news publisher, the lawsuit is more evidence that tech companies can’t be trusted with content. And that conclusion is correct. Ever since OpenAI’s then-CTO Mira Murati was caught like a deer in the headlights when asked about what training data had been used to train Sora, the company’s now-discontinued video model, it’s been clear AI companies will always take the most liberal view of “fair use” when it comes to harvesting content for their models.

    {“blockType”:”mv-promo-block”,”data”:{“imageDesktopUrl”:”https://images.fastcompany.com/image/upload/f_webp,q_auto,c_fit/wp-cms-2/2025/03/media-copilot.png”,”imageMobileUrl”:”https://images.fastcompany.com/image/upload/f_webp,q_auto,c_fit/wp-cms-2/2025/03/fe289316-bc4f-44ef-96bf-148b3d8578c1_1440x1440.png”,”eyebrow”:””,”headline”:”u003Cstrongu003ESubscribe to The Media Copilotu003C/strongu003E”,”dek”:”Want more about how AI is changing media? Never miss an update from Pete Pachal by signing up for The Media Copilot. To learn more visit u003Ca href=u0022https://mediacopilot.substack.com/u0022u003Emediacopilot.substack.comu003C/au003E”,”subhed”:””,”description”:””,”ctaText”:”SIGN UP”,”ctaUrl”:”https://mediacopilot.substack.com/”,”theme”:{“bg”:”#f5f5f5″,”text”:”#000000″,”eyebrow”:”#9aa2aa”,”subhed”:”#ffffff”,”buttonBg”:”#000000″,”buttonHoverBg”:”#3b3f46″,”buttonText”:”#ffffff”},”imageDesktopId”:91453847,”imageMobileId”:91453848,”shareable”:false,”slug”:””,”wpCssClasses”:””}}

    The all-or-nothing trap

    The problem with “don’t trust them” is that it often fuels a binary perspective: that the only reasonable reaction is to lock down your corpus, blocking AI bots from ever ingesting a single character. Opening it up, even a little, to a bot—even a supposedly “legit” one—means trusting the company to play by certain rules, a key one being: content used for AI search won’t be thrown into training data. So you can open up and hope, or block and stay safe.

    I see this more and more in my consulting work: the instinct to protect IP makes publishers reluctant to even do generative engine optimization (GEO) testing. This perspective is understandable, but ultimately self-defeating. Blocking bots means sacrificing visibility in AI answers. And while translating that visibility into good business outcomes is far from guaranteed, AI experiences are rapidly becoming the future. What publishers need is an approach that preserves AI as a path for audiences to discover them and build their authority while not taking on faith that the AI companies will play by the rules.
    Bot blocking is generally centered around the Robots Exclusion Protocol (aka robots.txt), which governs which bots can scrape content on a site. Importantly, there are different kinds of bots. For this discussion, you really only need to know that there are training bots and retrieval (i.e., search/answer) bots. The former harvests information into vast archives to build new models, while the latter grabs specific info to answer individual queries in real time. Training bots copy data and keep it; retrieval bots use it once, then poof. (The search bots that power discovery do keep an index, the way Google always has, but that’s a card catalog, not a model.)

    It’s becoming more or less standard for publishers to block training bots absent some kind of licensing deal. Retrieval, however, is how articles appear in AI answer engines. If your article is blocked, the engine only has metadata to go on, so if a competing site is open and yours is blocked, there’s a high likelihood the engine will favor your competition in the answer.

    This is where many publishers trip up. They want to compete in the answer, but they don’t trust the AI company to simply scrape the article for just the one query. Many will assume the AI company will keep that article and use it for either training or for allowing their users to access the full text—which might hurt even more if there’s a paywall. So they block everything since it’s better than risking giving away the store.

    Verification > trust

    There is a happy medium here. You can open up content to retrieval bots without blindly trusting the AI companies’ claims that they’ll never train on it. The approach starts with blocking training bots (of course) and then selectively allowing retrieval bots where AI visibility is important. Then, you build your own verification: Your CDN (content distribution network) checks every bot and verifies it against the vendor. It also logs the visit—what the bot scraped and when. That’s evidence you can use later if the vendor does something they shouldn’t.

    To monitor whether an AI vendor may be training on your retrieved content, you can seed your site with “tracer” phrases, checking whether they show up in the raw model with a set of queries run on a schedule. If they do, it’s a strong indicator that your content is being thrown into training data. And any testing that opens up content to new kinds of crawlers should be done in pieces—a slice of the content—before deploying site-wide. Problems? Reverse course with a single file change. And measure against the only outcome that matters: whether your visibility in AI answers is improving or not.

    While the media industry has good reason to be paranoid, it’s worth pointing out that the AI companies are incentivized to ensure their bots play by the rules. The companies that determine AI visibility (OpenAI, Anthropic, Perplexity, et al.) are the same ones spending fortunes on licensing deals and courtroom settlements. They’ve also learned, expensively, what courts do with sloppy acquisition. If they were to cheat and use retrieval content for training, that would transform a murky fair-use fight into clear evidence of misrepresentation. You don’t have to believe in their virtue to understand it’s in their interest to avoid that kind of exposure.

    A crawl is really a skim

    AI retrieval bots don’t actually read all your content anyway. When crawlers scan a web page, they expend the absolute minimum number of tokens to figure out what’s on the page to make a judgment about whether it’s worth citing. So even if you do make the entirety of a page available to crawlers, they often don’t read it—at least when it comes to determining presence in AI answers.

    But making the full text available to crawlers means there’s a greater chance of that extra context mattering in deeper, research-oriented queries—the exact kind where people tend to check their sources. Limiting what the crawlers can see means transferring your authority to more visible publishers for no real benefit.

    No one is asking publishers to trust companies that torrent online libraries. But the web has never run on trust—it runs on logs, verification, and who has the leverage. “Don’t trust” and “be discovered” aren’t opposites. Handled right, the first is how you afford the second.

    {“blockType”:”mv-promo-block”,”data”:{“imageDesktopUrl”:”https://images.fastcompany.com/image/upload/f_webp,q_auto,c_fit/wp-cms-2/2025/03/media-copilot.png”,”imageMobileUrl”:”https://images.fastcompany.com/image/upload/f_webp,q_auto,c_fit/wp-cms-2/2025/03/fe289316-bc4f-44ef-96bf-148b3d8578c1_1440x1440.png”,”eyebrow”:””,”headline”:”u003Cstrongu003ESubscribe to The Media Copilotu003C/strongu003E”,”dek”:”Want more about how AI is changing media? Never miss an update from Pete Pachal by signing up for The Media Copilot. To learn more visit u003Ca href=u0022https://mediacopilot.substack.com/u0022u003Emediacopilot.substack.comu003C/au003E”,”subhed”:””,”description”:””,”ctaText”:”SIGN UP”,”ctaUrl”:”https://mediacopilot.substack.com/”,”theme”:{“bg”:”#f5f5f5″,”text”:”#000000″,”eyebrow”:”#9aa2aa”,”subhed”:”#ffffff”,”buttonBg”:”#000000″,”buttonHoverBg”:”#3b3f46″,”buttonText”:”#ffffff”},”imageDesktopId”:91453847,”imageMobileId”:91453848,”shareable”:false,”slug”:””,”wpCssClasses”:””}}



    Source link

    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

    Related Posts

    What happened when Meta tried to collect employee data

    September 12, 2026

    The iPhone Duo isn’t just a folding phone

    September 12, 2026

    I bought a Birkin as an investment. It taught me a hard lesson about human judgment

    September 12, 2026
    Top News
    Business 4 Mins Read

    120,000 people applied for this very NSFW ‘hottest vacancy in AI right now’

    Business 4 Mins Read

    One AI company’s latest opening for a consultant role is anything but a standard tech…

    IHSA Football Preseason Poll Released

    August 30, 2026

    5 small shifts to turn creativity into a daily wellness practice

    March 26, 2026

    Why employees with chronic pain feel shame—and how they can break free

    March 29, 2026
    Top Trending
    Business 4 Mins Read

    What happened when Meta tried to collect employee data

    Business 4 Mins Read

    When employees at Meta learned that the company would start collecting their…

    World Politics 1 Min Read

    Georgia High School Pulls Off Stunning Last-Second Win

    World Politics 1 Min Read

    A Georgia high school football team delivered a stunning victory with a…

    Business 7 Mins Read

    Don’t trust AI companies with your content. Verify them instead

    Business 7 Mins Read

    If it’s a day ending in Y, you can count on a…

    Categories
    • Business
    • Economy
    • Headline News
    • Top News
    • US Politics
    • World Politics
    About us

    The Populist Bulletin was founded with a fervent commitment to inform, inspire, empower and spark meaningful conversations about the economy, business, politics, government accountability, globalization, and the preservation of American cultural heritage.

    We are devoted to delivering straightforward, unfiltered, compelling, relatable stories that resonate with the majority of the American public, while boldly challenging false mainstream narratives that seem to only serve entrenched elitists, and foreign interests.

    Top Picks

    What happened when Meta tried to collect employee data

    September 12, 2026

    Georgia High School Pulls Off Stunning Last-Second Win

    September 12, 2026

    Don’t trust AI companies with your content. Verify them instead

    September 12, 2026
    Categories
    • Business
    • Economy
    • Headline News
    • Top News
    • US Politics
    • World Politics
    Copyright © 2025 Populist Bulletin. All Rights Reserved.

    Type above and press Enter to search. Press Esc to cancel.