Skip to main content
frankvitetta.com

Projects / Open source tool

AI Scraper: collect AI answers
and their citations.

An open source tool that samples what Perplexity and ChatGPT answer and which sources they cite.

AI Scraper runs a list of prompts through Perplexity or ChatGPT in a real Chrome browser. It saves each answer together with the sources it cited.

AI Scraper on GitHub

Open source under the MIT licence. Please read the caution below before you use it.

What it is and who made it

AI Scraper is my own open source project. I made it at LLM Scout, which I founded. It is written in Python and uses Selenium to drive Chrome. It is released under the MIT licence. The code is public on GitHub.

Because it is my own project, please read this page as the maker describing his tool. It is not an independent review.

Read this first

What it is useful for

Prompt sampling means you fix a set of prompts and run each one more than once. You then read the results as an estimate. AI Scraper is one way to collect a small sample of that kind. For the prompts you choose, it records what Perplexity and ChatGPT answered and which sources each answer engine gave as a citation.

The limits of sampling still apply. A single run of a prompt is one sample. It is not the answer. If you change the prompts, the results change with them.

How it works

  1. You put your prompts in a CSV file, one prompt per row.
  2. The tool opens Chrome and enters the prompts one at a time, on Perplexity or on ChatGPT. It drives a real browser and paces its requests.
  3. It saves each prompt, the answer text and the addresses of the cited sources to a CSV file. JSONL output is optional.

ChatGPT needs you to log in yourself, in the browser window the tool opens. You can set a limit on the number of prompts, so a first run can be a small sample.

What you need

  • Python 3.
  • Google Chrome or Chromium.
  • The Python libraries the project depends on, including Selenium.
  • Your own ChatGPT account, if you want to sample ChatGPT.

The installation steps are in the README on GitHub.

Limits

  • It reads pages that were built for people, so it may break when either site changes its markup.
  • It covers Perplexity and ChatGPT only.
  • What it collects is a sample from one browser at one moment. It is not a measure of how often a brand is mentioned overall.
  • It is an independent project with no connection to Perplexity or OpenAI. Under the MIT licence it comes without warranty.

Get the code

The repository holds the source, the README and the licence.

View the repository on GitHub

github.com/frankmedia/ai-scraper