Every month, a non-profit called Common Crawl fetches over 2 billion webpages and gives the copy away to anyone who wants it. That free archive is where most AI training data starts. When Mozilla audited the LLMs released between 2019 and 2023, 64% had trained on Common Crawl data in some form, and GPT-3 was […]
How To Remove Negative Content From Google Before AI Answers Cite It
A client sends you a link. A news story, a mugshot site, a people-search page with their home address on it, a Reddit thread calling the business a scam. The question is always the same: can you make this go away? The answer depends entirely on what kind of page it is, and the most […]