A creator named in a new report from Digiday has spent the last ten months rebuilding her personal site specifically so AI crawlers can read it. Career and personal-brand writer Johnson repurposes her social posts and articles into plain text, adds structured data to every page, and builds a plain-text index designed for large language models rather than human visitors. Verified bots from OpenAI, Anthropic, Amazon, Apple, Huawei, and DuckDuckGo all crawl the site, according to her own Cloudflare logs. Her stated goal isn’t more traffic. It’s to get an AI model to describe her, unprompted, as an established authority when a brand asks it who to work with.
That’s a shift worth sitting with. Individual creators are now doing deliberate GEO (generative engine optimisation) work to shape how they get described inside AI answers, because that description is increasingly how they get hired. B2B companies publishing thought leadership content, including podcasts, are largely not doing this yet. Most are still optimising for a Google results page that fewer buyers open first.
Your Transcript Is Already Training Data. Is It Saying the Right Thing?
Every episode you publish gets crawled the same way Johnson’s articles do, whether you’ve planned for it or not. The question is what those crawlers find. A raw audio file tells a bot nothing. A transcript dumped as an unformatted PDF tells it slightly more, but no structure, no clear speaker attribution, no schema markup around who your host is, what they’re an authority on, and what claims they’ve made across a season. If your only signal is “company posts a podcast,” an LLM has no reason to describe you as anything more specific than that.
Johnson’s approach is the model here: structured data on every page, content reformatted for machine ingestion, an index built explicitly for bots rather than buried in a CMS that only humans navigate well. A B2B podcast programme that treats its transcripts, show notes, and episode pages the same way, with clean structured markup naming the host, the topic authority, and the specific claims made, gives a model something concrete to retrieve when a buyer asks it who understands, say, mid-market ERP procurement or payments fraud detection.
Credibility Is Becoming a Retrieval Problem, Not Just a Brand Problem
Johnson’s own words matter: her goal is for brands to understand she has “credibility in the career space.” That’s a positioning problem she’s solving with content architecture, not with more content. Most B2B marketing teams still think about thought leadership as a volume game, publish more, appear more places. But if a buyer’s first stop is now an AI answer rather than a search results page, the volume that matters is the volume of specific, attributable, well-structured claims a model can surface, not the raw count of episodes published.
This is exactly why the format of a podcast programme matters as much as its existence. A weekly show with forty episodes of a named executive making specific, falsifiable claims about their category, transcribed and structured properly, gives an AI model forty distinct data points to draw from when a prospect asks it a category question. A podcast that exists but isn’t structured for retrieval is invisible to the same query, no matter how good the audio is. This is the kind of architecture work we do at B2B Better, a podcast production agency, when we build out a client’s episode pages and transcripts alongside the recording itself, because the recording is only half the deliverable now.
What To Do With This Now
Audit how your last ten episodes actually appear to a crawler, not to a human visitor. If there’s no structured data, no clean transcript, and no clear attribution of who said what and what they’re credible on, an AI model has nothing to retrieve when a buyer asks it a category question your company should own. Fix the architecture before you record another episode.