There is no single, established tool called "GEO CLI". The name is used for unrelated projects: the GitHub repository tamnd/geo-cli, for example, is a Go command line for public geo data and has nothing to do with Generative Engine Optimization. If you mean GEO in the sense of AI search visibility, a "GEO CLI" is a set of command-line checks that tell you whether AI crawlers can reach and read your pages: server responses per bot, robots.txt, content in the raw HTML, JSON-LD and llms.txt. This guide shows each check with curl and how to run all of them with the open-source geo-tool-check.
An earlier version of this article described "GEO-CLI" as a marketing framework. That framework was not an established method, so we replaced it with what you can actually run and verify.
What "GEO CLI" can mean
| Name | What it is | Related to AI search visibility? |
|---|---|---|
| tamnd/geo-cli | Go command line that reads public geo data over HTTPS; the README calls it a fresh scaffold. Apache 2.0 license. | No |
AWS CLI geo-places, geo-maps | Commands of the AWS CLI for Amazon Location Service (places, maps) | No |
| GEO audit CLIs on GitHub | Several small open-source projects that score pages for AI search readiness. Maturity varies; check license, last commit and what exactly is measured before you rely on one. | Yes |
| geo-tool-check | Our own CLI and MCP server, published on npm, MIT license, runs locally | Yes |
GEO in this article means Generative Engine Optimization, the term introduced by Aggarwal et al. in the paper "GEO: Generative Engine Optimization" (2023) for making content more visible in answers of generative search engines.
Why the command line is useful for GEO
Most reasons a page never shows up in ChatGPT search, Claude or Perplexity are technical and binary: the bot gets a 403, robots.txt disallows it, or the text only exists after JavaScript runs. A browser hides all of that. A command line shows you what a plain HTTP client receives, is quick to repeat after every deploy and fits into CI. What it cannot show is whether an AI assistant actually cites you. That needs prompt monitoring, see the last section.
All examples use https://www.example.com/. Replace it with your own URL.
Check 1: Does each AI bot get a 200?
Firewalls and CDN bot rules often match on the user-agent. This loop requests the same page with the tokens of the main AI bots and prints the status code:
for ua in GPTBot OAI-SearchBot ChatGPT-User ClaudeBot Claude-SearchBot PerplexityBot; do
printf "%-18s" "$ua"
curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 (compatible; $ua)" https://www.example.com/
done
A 403, 429 or a redirect to a challenge page for one bot while a normal browser gets 200 points to a bot rule. Keep in mind what this test can and cannot do: it simulates user-agent based rules only. Many bot protections also check whether the request really comes from the operator's IP ranges, which you cannot fake from your machine. OpenAI, for example, publishes the IP ranges of its bots on its crawler overview.
Check 2: What does robots.txt say?
# status and content type: should be 200 and text/plain
curl -s -o /dev/null -w "%{http_code} %{content_type}\n" https://www.example.com/robots.txt
# show the groups for specific AI bots
curl -s https://www.example.com/robots.txt | grep -iA4 -E "user-agent: *(gptbot|oai-searchbot|claudebot|claude-searchbot|perplexitybot|\*)"
Under RFC 9309, a crawler follows the group that matches its user-agent most specifically and then ignores the generic * group. A rule under User-agent: * therefore does not apply to a bot that has its own group. Which bots you should allow for AI search and which ones only collect training data is covered in 7 rules for robots.txt: AI bots to allow.
Check 3: Is the content in the raw HTML?
Take a sentence that is visible on your page and search for it in the unrendered HTML:
curl -s https://www.example.com/ | grep -c "a sentence from your page"
# rough word count of the raw HTML without tags
curl -s https://www.example.com/ | sed -e 's/<script.*<\/script>//g' -e 's/<[^>]*>//g' | wc -w
A result of 0, or a word count far below what you see in the browser, means the content is loaded by JavaScript. Crawlers that do not render JavaScript then see an almost empty page. How to fix that is explained in making JavaScript sites accessible to AI crawlers.
Check 4: Is the JSON-LD valid?
This one-liner extracts every JSON-LD block and parses it. Invalid JSON fails loudly, which is exactly what you want:
curl -s https://www.example.com/ | python3 -c '
import sys, re, json
html = sys.stdin.read()
blocks = re.findall(r"<script[^>]*application/ld\+json[^>]*>(.*?)</script>", html, re.S)
print(len(blocks), "JSON-LD block(s)")
for b in blocks:
data = json.loads(b)
print(data.get("@type") if isinstance(data, dict) else "list/graph")
'
Valid JSON is only the first step. Whether the markup qualifies for rich results in Google is checked with Google's Rich Results Test.
Check 5: Is there an llms.txt?
curl -s -o /dev/null -w "%{http_code} %{content_type}\n" https://www.example.com/llms.txt
curl -s https://www.example.com/llms.txt | head -20
llms.txt is a proposal for a Markdown file that gives language models a curated overview of a site. It is not a standard and does not control access. Treat it as optional; the robots.txt and HTML checks above matter more. Details on structure and use are in our guide to llms.txt and Markdown for AI agents.
All checks in one command: geo-tool-check
If you do not want to maintain your own scripts, geo-tool-check bundles these checks. It is published on npm under the MIT license, needs Node.js 18 or newer, no account and no API key, and runs entirely on your machine.
# full check: score 0 to 100 in six categories
npx geo-tool-check example.com
# only robots.txt: which AI crawlers are blocked
npx geo-tool-check example.com --crawlers
# machine-readable output
npx geo-tool-check example.com --json
# German labels
npx geo-tool-check example.com --lang de
Node does not execute JavaScript, so the tool reports what a plain HTML fetch returns. A client-rendered page scores low, and that is the finding. The same package also runs as an MCP server (npx -y geo-tool-check --mcp), so an AI agent in your editor can run the checks while you work.
Run the checks in CI
The most expensive GEO mistake is a deploy that blocks AI bots by accident, for example a new robots.txt from a staging template or a stricter firewall rule. The --min-score flag sets exit code 1 below a threshold, so a pipeline step can catch it:
npx geo-tool-check https://www.example.com/pricing --min-score 70
Run it against your most important pages after each production deploy. In the same step, the curl loop from check 1 catches bot rules that return a 403.
What a CLI cannot tell you
All checks above answer one question: can AI systems read your page? None of them tells you whether ChatGPT, Perplexity or Gemini actually mention or cite your brand for the questions your customers ask. For that you need to run real prompts regularly and record the answers and sources. Our AI visibility checker does this for your own set of questions; pricing is on request.
FAQ
Is GEO-CLI a real tool?
There is no single established tool of that name. The GitHub repository tamnd/geo-cli is a command line for geo data, not for Generative Engine Optimization. Several small open-source projects offer command-line audits for AI search readiness, among them our geo-tool-check.
How do I check from the command line whether GPTBot can access my site?
Request the page with curl and a user-agent that contains GPTBot, for example curl -s -o /dev/null -w "%{http_code}" -A "Mozilla/5.0 (compatible; GPTBot)" https://www.example.com/, and check that the status is 200. Then check robots.txt for a GPTBot group or a matching rule under User-agent: *.
Does curl show exactly what AI crawlers see?
Close, but not exactly. curl shows the raw HTML without JavaScript, which is what crawlers without rendering receive. It cannot reproduce IP-based bot verification, so a CDN may treat the real crawler differently from your test request.
Do I need an llms.txt?
No. llms.txt is a proposal, not a standard, and it does not control crawler access. It can be a useful extra for documentation-heavy sites, but robots.txt, server responses and server-rendered content are more important.
Can I run GEO checks in a CI pipeline?
Yes. npx geo-tool-check with the --min-score flag exits with code 1 when a page scores below the threshold, and the --crawlers flag checks only robots.txt. Both work as a step after each deploy.
Ready for better AI visibility?
Test now for free how well your website is optimized for AI search engines.
Start Free AnalysisRelated GEO Topics
Share Article
About the Author
- Structured data for AI crawlers
- Include clear facts & statistics
- Formulate quotable snippets
- Integrate FAQ sections
- Demonstrate expertise & authority


