SEO + GEO
Optimize pages for both classic search retrieval and AI citation. SEO is the discovery layer; GEO is the extraction, entity understanding, and recommendation layer.
The 5 Required Rules
These rules are the minimum routing checklist encoded in the SEO/GEO skills, not an exhaustive SEO course. They keep agent work focused on page structure, entities, schema, and crawlability before deeper growth work. Source: seo-geo/SKILL.md; seo-geo/references/seo-geo-playbook.md
- Keyword placement -- primary keyword in at least 4 of 6 locations: URL slug, title tag, meta/social title, H1, at least one H2, body copy (ideally intro).
- Entity placement -- primary entity and supporting entities in title, H1, intro, core sections, and schema.
- Entity gap analysis -- extract entities from top competitors, add missing relevant entities that improve topical completeness. Do not add irrelevant brands.
- JSON-LD schema -- place in page
<head>, not only through client-side injection. - Schema entity reinforcement -- mirror on-page entities inside schema fields (
name,description,about,sameAs,knowsAbout,mentions,author,publisher).
Jesper Nissen's screenshot is a small field proof for why these rules stay bundled. The root post reports a homepage ranking jump after keyword placement, entity placement, entity-gap completion, and schema/entity reinforcement; the attached reply claims a CTR bump after applying the entity-gap tip and says schema is still underused. Treat this as anecdotal operational evidence for the existing checklist, not as a new ranking guarantee. Source: X/@JespernissenSEO and local image review, 2026-07-04
Master GEO Scorecard
| Category | Weight |
|---|---|
| AI citability and extractability | 25% |
| Brand authority and off-site signals | 20% |
| Content quality and E-E-A-T | 20% |
| Technical foundations | 15% |
| Structured data | 10% |
| Platform optimization | 10% |
Bands: 90-100 elite, 75-89 strong, 60-74 mixed, 40-59 weak, 0-39 critical.
AI Citability Engineering
Models cite passages that are direct, self-contained, fact-rich, structurally legible, and make sense outside their original context. The 134-167 word passage target is a heuristic for high-value answer blocks.
5-part citability model: (1) Answer block quality -- question-based H2, direct answer in first 1-2 sentences, named entities, precise claims. (2) Self-containment -- names subject explicitly, no dangling pronouns, at least one concrete fact. (3) Structural readability -- one H1, clean H2-H3 nesting, short paragraphs, lists for steps, tables for comparisons. (4) Statistical density -- percentages, dollar values, dated claims, named studies, benchmarks. (5) Uniqueness -- original data, case studies, experiments, screenshots.
E-E-A-T
Experience: first-hand language, implementation notes, screenshots, case studies with numbers. Expertise: visible author, credentials, sourced claims, correct terminology. Authoritativeness: external citations, strong topic coverage across the site, media mentions. Trustworthiness: clear contact, privacy policy, accurate claims, visible dates, HTTPS.
Technical GEO Foundations
Crawlability: choose discovery, user-fetch, and training permissions separately. Google Search AI features use Googlebot and Search preview controls; Google-Extended affects specified Gemini training/grounding uses and has no effect on Google Search. ChatGPT search uses OAI-SearchBot while GPTBot is the separate possible-training control. Claude search uses Claude-SearchBot/Claude-User while ClaudeBot is the training crawler. Perplexity documents PerplexityBot/Perplexity-User as non-training search/fetch routes. Verify robots policy and CDN/WAF/server receipts before declaring a surface crawlable. Source: Google, OpenAI, Anthropic, and Perplexity first-party crawler documentation, 2026-08-11 review
Rendering: prefer SSR/SSG. Make schema server-rendered. Some AI crawlers are limited in JavaScript execution.
Core Web Vitals: LCP < 2.5s, INP < 200ms, CLS < 0.1.
Internal linking: high-value pages within ~3 clicks of homepage, descriptive anchor text.
Googlebot byte ceiling: Googlebot currently fetches up to 2MB for any individual URL, excluding PDFs, and that counter includes the HTTP header. If the HTML response exceeds that limit, Google processes the first 2MB as the fetched resource; referenced CSS/JS/image/video resources have their own fetch limits, and PDFs use a 64MB limit. Keep title, robots directives, canonicals, structured data, and primary content before bloated inline assets, menus, or hydration payloads. Source: Google Search Central, 2026-03-31; Source: local artifact review, 2026-07-03
Schema Design
Baseline graph: Organization/Person + WebSite + page-specific type + BreadcrumbList. Use sameAs to disambiguate across Wikipedia, Wikidata, LinkedIn, YouTube, GitHub, Crunchbase, etc. Prefer one @graph block with @id references. Use absolute URLs and ISO dates.
llms.txt
Optional machine-readable summary at domain root: what the site is, what sections matter, canonical pages, and core facts. Keep it aligned with the real site and test whether target agents consume it. It is not required for Google Search AI features, a universal ranking signal, or a replacement for useful content, ordinary crawlability, or semantic structure.
Competitive Comparison Capture
Capture branded commercial-intent queries directly. Publish exact-match vs, review, and alternatives pages. Make answer blocks so explicit that AI Overviews can lift them verbatim.
Per competitor: at minimum one standalone review page + one exact-match comparison page. Structure for answer extraction: question-led H2s, quotable verdict sentences, side-by-side tables, FAQ blocks.
Comparison page anatomy: exact-match title/slug/H1, first-paragraph verdict, Key Takeaways block, side-by-side table, pricing with dates, best-for guidance, honest tradeoffs, FAQ, clear CTA.
Firsthand testing required: sign up, use the products, take notes and screenshots. Fake reviews are obvious. Combine firsthand testing + external review language (G2, Capterra, Reddit) + your own positioning.
AIO conquest loop: pick 3-6 named competitors, ship one honest competitor-vs-product page and one standalone review per competitor, then inspect Google AI Overview / AI Mode / ChatGPT / Perplexity answers for those exact queries. The durable pattern from Kleo's bookmark is not "write attack pages"; it is firsthand use + external review language + extraction-friendly verdicts. Their screenshot shows Google returning a direct AI Overview answer for "kleo vs stanley" and citing Kleo's comparison/review pages. Source: X/@RobHoffman_, 2026-03-13
Local-service agent loop: for local businesses, route Claude/Cowork/browser agents through competitor page extraction, high-intent local keyword discovery, Google Business Profile completeness, service-page locality signals, review-response drafting, schema checks, and weekly rank/citation monitoring. The Claude Cowork bookmark is useful as a workflow sketch, but its attached images are only Claude and Google Business Profile logos; the durable asset is the checklist, not the visuals. Source: X/@bloggersarvesh, 2026-03-14
Programmatic SEO Domain Boundary
Programmatic SEO is a brand-authority decision, not only a template-generation tactic. Thin, short-term, or unproven pSEO experiments should not live on the primary brand domain: if they collapse, they can dilute topical trust, create crawl waste, and leave the brand with a footprint of disposable pages. Main-domain pSEO is reserved for durable pages with unique value, first-party or defensible data, useful internal links, clear indexation rules, and a long-term maintenance owner. Disposable experiments belong on a separate property, in noindex/staging, or outside the brand graph until quality and demand are proven.
The EXM row is useful as a blunt risk boundary rather than a precise traffic forecast: do not borrow the main site's authority for mass pages that are meant to be short-lived. Build authority and backlinks for the durable domain; keep churny pSEO tests isolated. Source: X/@EXM7777, 2026-03-12
Programmatic generation is not index admission
Jake Ward's resolved Byword article contributes a sound software boundary:
versioned niche taxonomy -> strict JSON payload -> validation -> page-type
renderer. It is better than freeform generated HTML because content, structure,
and presentation can be tested and replaced independently. Its two utility
questions are equally durable: would the page help without search engines, and
would someone intentionally return to it? Source: X Article via
raw/x-bookmarks/enriched/2031701565434732917.json, 2026-03-11
The reported outcome is not durable proof. The article self-reports 13,000+
pages, weekly clicks rising from 971 to 5,500 in 60 days, roughly half the
pages indexed, and no negative Google signal. By 2026-08-11, Byword's own
robots.txt said /resources/ had been retired during a Google “scaled
content abuse” manual-action recovery and left crawlable so Googlebot could
observe 410 Gone. Two matching resource examples returned 410; all 13
current sitemaps contained 863 URL entries. The current count does not recreate
the earlier cohort, but it directly contradicts using the 60-day snapshot as a
safety or permanence claim. Source: Byword robots.txt,
Byword sitemap, and captured HTTP responses,
2026-08-11
The correct system separates three decisions:
generate candidate -> validate payload -> render and inspect
-> stage/noindex cohort -> admit selected URLs to index
-> monitor -> refresh, consolidate, noindex, or remove
Index admission needs page-level utility, source and rights provenance, factual checks, batch similarity and cannibalization checks, raw HTTP metadata and content, rendered behavior, accessibility/performance, an owner, an index budget, and a removal/repair path. Type-valid JSON proves shape, not value. Purpose-built UI can improve value, but an interactive checkbox or filter does not rescue generic or unsupported content.
Google's current policy is purpose- and value-based: many automatically or manually generated pages can be scaled content abuse when their primary purpose is ranking manipulation rather than helping users. For a manual action, fix all affected patterns, keep the relevant URLs reachable for review, document the quality issue and remediation, and request reconsideration only after the systemic cause is corrected. Source: Google scaled-content policy, generative-AI guidance, and Manual Actions report, reviewed 2026-08-11
Search Console Query Refresh Loop
Google Search Console is the cheapest refresh signal because it shows the queries a page already earns impressions for. Use it for pages that have traction but under-capture intent.
Workflow:
- Pick one page with meaningful clicks or impressions in the last 90 days.
- Export the page-filtered query CSV from Search Console.
- Ask Claude to group queries into satisfied intent, missing intent, title/H2 opportunities, FAQ opportunities, and internal-link opportunities.
- Update the page with answer blocks, missing subtopics, schema fields, and clearer headings.
- Re-measure clicks, impressions, CTR, and average position after the next crawl window.
The reviewed artifact shows the intended input shape: a Search Console performance view with hundreds of clicks, six-figure impressions, low CTR, and average position around page one. The durable lesson is not the screenshot metrics; it is the page-refresh loop from query evidence to editorial changes. Source: X/@ayushtweetshere, 2026-03-14
The reviewed @hridoyreh artifact sharpens the target selection: filter Search Console for high-impression queries where an existing page ranks around positions 8-20, then update that page with missing answer blocks, People Also Ask coverage, headings that match search intent, and a few internal links from stronger pages. This is a refresh loop for pages already near the prize, not a generic "publish more posts" tactic. Source: X/@hridoyreh and local image review, 2026-07-03
AI Search Measurement Loop
Treat AI-search visibility as a parallel measurement lane, not a replacement label for SEO. The primary John Mueller reply says businesses that earn referred traffic should inspect the full picture, actual usage metrics, and audience percentages before allocating time. It does not say Google ranking determines inclusion in ChatGPT, Claude, Gemini, or Perplexity, and it does not endorse the attached vendor packages. Source: John Mueller primary Reddit comment
Google's June 2026 Search Console release changes the operating loop. A subset of properties now has dedicated Search and Discover generative-AI reports with impressions, pages, countries, devices where supported, and time granularity. The data remains in the overall report. Use the dedicated report when available; retain overall Web/Performance and first-party conversion analytics as the baseline. Google explicitly warns that no third-party tool has access to its internal ranking or AI systems. Source: Google Search Console announcement; Source: Google generative-AI optimization guide
| Evidence lane | Record | What it proves |
|---|---|---|
| Audience allocation | Channel usage %, target segment, business value, time/cost allocation | Whether the surface deserves work |
| Google Search AI | Dedicated Search Console report when available; overall report fallback | Google-reported impressions/pages/time trends, not another engine's visibility |
| Answer observation | Exact query, product/model, date/time, locale, account state, mode, mention, cited URL, screenshot/export | What one controlled run displayed |
| Stability | Repeated trials over time with the same panel | Volatility and directional change, not causality by itself |
| Business outcome | First-party referral, engagement, signup, revenue, or other conversion event | Whether observed visibility created value |
| Third-party tracker | Disclosed query panel, collection method, timestamps, exports | A vendor's reproducible observation—not internal rank or causal truth |
The complete @alexgroberman replay reviewed all five SEO Stuff links and four images. The Search Console screenshot lacks property/date/filter context; six growth charts lack intervention and comparison data; one ChatGPT 5.2 answer is anecdotal; and the citation-column screenshot is a useful monitoring-interface pattern. Keep the offer as market signal and the interface as design signal. Hold the 60–90-day results, client counts, backlink causality, “safe to repeat,” and “single biggest factor” claims until independently reproducible. Source: X/@alexgroberman; local replay receipt: .brain/artifacts/x/2032084993368080515/seo-geo-audience/source-snapshots.md.
Platform-Specific Notes
Google Search AI features: ordinary Googlebot index/snippet eligibility, valuable non-commodity content, semantic structure, truthful structured data, and Search Console measurement. ChatGPT: verify OAI-SearchBot, keep important facts public and stable, and use a repeated prompt/referral panel. Perplexity: verify PerplexityBot and record the exact cited pages over repeated trials. Claude: verify Claude-SearchBot/Claude-User; decide ClaudeBot training access separately. Gemini: do not infer Google Search inclusion from Google-Extended or a Gemini answer.
Audit Workflow
- Define target (keyword, entity, page type, geography, platform mix)
- Fast SEO/GEO audit (title, H1, slug, intro, schema, canonical, dates, author, crawlability)
- Keyword and entity gap
- Score (citability, E-E-A-T, technical, schema, platform fit)
- Rewrite (title, H1, H2 outline, intro, answer blocks, schema)
- Strengthen graph (author pages,
sameAs, external profiles,llms.txt) - Deliver prioritized action plan
Agent Skills (Enforced)
Route SEO/GEO work through the skill registry rather than hard-coded home-directory projections:
ai-seo(skills/personal/ai-seo/SKILL.md) — AI-search visibility, source/crawler policy, repeated prompt panels, and measurement.programmatic-seo(skills/personal/programmatic-seo/SKILL.md) — page-family strategy, typed generation, candidate/index separation, staged cohorts, and scaled-content recovery.seo-audit— traditional crawlability, indexation, technical, on-page, content, and authority audit.schema— accurate structured-data implementation when the visible content supports it.og-metadata-audit— OpenGraph/Twitter card metadata and preview behavior.
Use wiki/meta/skill-registry.md and wiki/SKILL-RESOLVER.md for the current installed paths. This page owns the knowledge model; executable behavior belongs in the skill source and must be regenerated into runtime projections after edits.
OG / Social Preview Lessons
Learned from Loop (loooop.dev) cross-referenced against Dedalus monorepo (April 2026):
twitter:namespace is permanent - X/Twitter did not change the meta tag namespace. The crawler still readstwitter:card,twitter:image, etc.- Next.js middleware naming is strict - file must be
middleware.ts, export must bemiddlewareordefault. A file namedproxy.tsis silently ignored. This was the root cause of broken Twitter cards on Loop. max-image-preview: "large"in rootrobotsmetadata tells Google (and helps Twitter) to render full-size image cards instead of thumbnails.- Canonical inheritance trap - In Next.js App Router, child pages that don't export
alternates.canonicalinherit the root layout's canonical (usually/). Every page needs its own. htmlLimitedBots+ middleware - Two complementary layers. Middleware intercepts social bots with custom HTML.htmlLimitedBotsinnext.config.tsserves lightweight HTML to SEO/AI crawlers. They overlap for social bots buthtmlLimitedBotsextends coverage.- DRY pattern -
buildPageMetadata(input)returns complete{ openGraph, twitter, alternates, title, description }from one function call. Eliminates 15-line boilerplate per page and ensures canonical consistency.
Twitterbot delivery invariant
When an owned project serves dynamic OG routes, its robots.txt or Next.js robots.ts should carry an explicit Twitterbot group for the shared aliases, even when User-agent: * already allows the whole site:
User-Agent: Twitterbot
Allow: /api/og
Allow: /api/og-home
Also allow the routes the project actually emits, such as /opengraph-image, /twitter-image, /og, or /og.png. This explicit group is defense in depth for projects whose wildcard rule blocks /api/; it is not proof that a social card will render and it does not invalidate a cached failed scrape. Diagnose the public delivery chain with a Twitterbot/1.0 user agent: the page must return 200, expose absolute og:image and twitter:image URLs, and the selected image URL must return 200 with an image content type. If those checks pass but the card is stale, deploy the metadata/image change and use a versioned image URL or new route to force a fresh scrape. In the July 27 Local Search audit, Twitterbot already received 200 under the wildcard allow rule while production still served the older metadata, proving that stale deployment or social cache—not robots alone—was the active failure mode. Source: User request; live Twitterbot delivery audit of local-search-xi.vercel.app, 2026-07-27
Tools
GEO-SEO-Claude (GitHub): open-source Claude Code GEO/AEO reference with /geo audit, citability, crawler, schema, brand, platform, report, and PDF routes. Keep it reference-only unless a future pass extracts a missing capability into Kevin's active ai-seo, seo-geo, seo-geo-optimization, or seo-audit skills. Source: GitHub/source review, 2026-07-03
Jam / SpreadJam: growth-channel agent platform that turns AI-search visibility and SEO/GEO audits into workflow actions. Use Kevin's installed SEO/GEO skills for direct audits; evaluate Jam when the task needs repeated growth-channel workflows, repo/site analysis, content-gap actions, warm leads, inboxes, or deliverability loops. Source: SpreadJam site/docs and local X video review, 2026-07-02
Cloudflare /crawl endpoint: single API call to crawl an entire site via Cloudflare Browser Rendering. Alternative to Firecrawl for automation use - compare latency, JavaScript rendering fidelity, and rate limits.
Signals (April 2026)
- Googlebot 2MB fetch ceiling (Google Search Central, March 2026): Googlebot currently fetches up to 2MB for an individual URL, excluding PDFs, and the counter includes HTTP headers. The SEO implication is response-ordering, not "total page weight": keep metadata, canonical tags, schema, and primary content early in the HTML response, while external resources are fetched separately with their own limits. PDFs get 64MB. Source: X/@Charles_SEO local screenshot review, 2026-07-03
- Claude Code SEO workflow (Cody Schneider, March 2026, 4K bookmarks): Set up a
.envwith Keywords Everywhere API key, DataForSEO API key, data warehouse for Google Search Console, and CMS access. Then use Claude Code for: keyword universe mapping (related keywords + SERP API for gap analysis), programmatic landing pages at scale (per-vertical with semantic terms + schema), link building (DataForSEO domain intersection for competitor backlink gaps), internal linking maps (topical relevance clusters from keyword data), and content refreshing. Source review resolved the self-thread to Graphed, a hosted marketing-agent service that positions this workflow as an autonomous growth operations system. Source: X/@codyschneider self-thread and Graphed site review, 2026-07-04 - Free backlinks from high-authority domains: WordPress and Medium both offer free backlink surfaces - republish or cross-post carefully and set canonical links back to the source page when appropriate. The reviewed WordPress artifact shows a high-authority WordPress domain snapshot (DR 95, AR 31), so route this as authority-surface evidence, not a license to spray thin duplicate pages. Source: X/@kalashvasaniya and local image review, 2026-07-03
- Google's John Mueller recommends tracking AI search traffic separately in analytics. No standard referrer header yet, but Perplexity and ChatGPT are increasingly identifiable.
- Awesome Design MD (GitHub): reverse-engineered design systems of Apple, Spotify, Airbnb, and other companies as single markdown files. Useful as reference data for schema and entity modeling on design-related pages.
Taxonomy
12 operational categories: Foundations, Research & Intent, Site Architecture, Technical Crawl & Indexing, On-Page Entities & Schema, Content Systems & Page Types, GEO/AEO/AI Platforms, Authority Links & Brand, Distribution & Feedback, Measurement & Experiments, International, Growth Systems. Cross-cutting: UX, conversion, accessibility, legal, performance, editorial quality.
Hridoy Rehman's folder-style SEO bookmark is useful as a curriculum-shaped view of the same taxonomy: foundations, research, SERP analysis, search intent, competitor research, and keyword gaps belong before tactical page edits. Because the enriched record has no resolved source beyond the tweet and is visibly truncated, it is supporting structure for this page rather than a separate SEO skill or a claim about a complete course outline. Source: X/@hridoyreh, 2026-03-15
Timeline
- 2026-08-11 | Reopened Jake Ward's formerly inaccessible X Article after full enrichment. Kept its taxonomy/schema/renderer architecture, but replaced the 60-day safety story with a candidate-to-index admission contract after Byword's current first-party
robots.txtdisclosed a scaled-content-abuse manual-action recovery and retired resources. Source: X2031701565434732917; Byword live site; current Google Search guidance - 2026-08-11 | Replayed all five links and four images from @alexgroberman against the primary John Mueller reply and current Google/OpenAI/Anthropic/Perplexity documentation. Replaced stale Search Console and crawler rules, split mention/citation/impression/referral/conversion evidence, and retained the vendor package only as market signal. Source: X/@alexgroberman; Source: Google Search Console report announcement
- 2026-07-27 | Established the cross-project Twitterbot delivery invariant: explicitly allow
/api/ogand/api/og-home, also allow each project's real image routes, verify the public page and image as Twitterbot, and treat versioned image URLs as the cache-busting path after deployment. Source: User request; live Twitterbot delivery audit of local-search-xi.vercel.app, 2026-07-27 - 2026-07-04 | Added the pSEO domain-boundary guardrail from EXM's text-only bookmark and mirrored it into the
programmatic-seoexecutable skill: main-domain pSEO requires durable unique value and maintenance; disposable/thin experiments stay isolated. Source: X/@EXM7777, 2026-03-12 - 2026-07-04 | Reviewed Jesper Nissen's entity-gap/schema artifact and kept it as reinforcement for the existing five-rule checklist: keyword placement, entity placement, entity gaps, JSON-LD, and schema entity reinforcement work together. Source: X/@JespernissenSEO, 2026-03-10
- 2026-07-04 | Reviewed @hridoyreh's folder-style SEO taxonomy bookmark as supporting structure for the existing playbook. No new SEO skill was created because the page and installed SEO/GEO skills already own the workflow. Source: X/@hridoyreh, 2026-03-15
- 2026-07-04 | Source-reviewed Cody Schneider's Claude Code SEO workflow and resolved the self-thread product link into Graphed. Kept local execution in the SEO/GEO skills; Graphed is tracked as a hosted workflow/vendor reference. Source: X/@codyschneider, 2026-03-02
- 2026-07-03 | Deep-reviewed two SEO screenshots: @hridoyreh's high-impression position 8-20 Search Console refresh tactic and @kalashvasaniya's WordPress high-authority backlink artifact. Source: X/@hridoyreh, 2026-03-12; Source: X/@kalashvasaniya, 2026-03-24
- 2026-07-03 | Deep-reviewed the @Charles_SEO Googlebot screenshot against the official Google Search Central blog post and corrected the 2MB rule from "total page weight" to "per-URL fetched response bytes, including headers." Source: X/@Charles_SEO, 2026-03-31; Source: Google Search Central, 2026-03-31
- 2026-07-03 | Deep-reviewed the @alexgroberman AI-search screenshots and reframed John Mueller's GEO/SEO reply as a measurement/audience-behavior signal. Added the separate AI-search measurement loop: Search Console + AI answer mentions/citations + AI Overview citation columns. Source: X/@alexgroberman, 2026-03-12; Source: local artifact review, 2026-07-03
- 2026-07-02 | Added the Search Console query-refresh loop from the Claude + GSC bookmark and reviewed the local GSC screenshot artifact. Source: X/@ayushtweetshere, 2026-03-14; Source: wiki/assets/x-bookmarks/2032834673433510012/image-01.jpg
- 2026-07-02 | Added Jam (SpreadJam) as a tool candidate for productized SEO/GEO-to-growth workflows after reviewing the @jia_seed Jam demo, current SpreadJam site/docs, and
wespreadjam/jam-nodessource snapshot. Source: X/@jia_seed, 2026-03-03; Source: SpreadJam, 2026-07-02 - 2026-06-30 | Deep-reviewed two high-priority SEO/GEO bookmark threads and local images. Added the competitor-comparison/AIO conquest loop, the local-service agent loop, and the requirement that comparison pages use firsthand product testing plus external review language rather than synthetic fluff. Source: X/@RobHoffman_, 2026-03-13; Source: X/@bloggersarvesh, 2026-03-14
- 2026-06-17 | Added inline provenance and normalized links to the canonical social content page. Source: seo-geo/SKILL.md; seo-geo/references/seo-geo-playbook.md
- 2026-04-07 | Added OG/social-preview lessons from Loop and Dedalus comparisons. Source: Loop and Dedalus monorepo review notes
- 2026-04-06 | Created SEO/GEO playbook page from skill references. Source: seo-geo/SKILL.md; seo-geo/references/taxonomy.md