SEO + GEO

Optimize pages for both classic search retrieval and AI citation. SEO is the discovery layer; GEO is the extraction, entity understanding, and recommendation layer.

The 5 Required Rules

These rules are the minimum routing checklist encoded in the SEO/GEO skills, not an exhaustive SEO course. They keep agent work focused on page structure, entities, schema, and crawlability before deeper growth work. Source: seo-geo/SKILL.md; seo-geo/references/seo-geo-playbook.md

  1. Keyword placement -- primary keyword in at least 4 of 6 locations: URL slug, title tag, meta/social title, H1, at least one H2, body copy (ideally intro).
  2. Entity placement -- primary entity and supporting entities in title, H1, intro, core sections, and schema.
  3. Entity gap analysis -- extract entities from top competitors, add missing relevant entities that improve topical completeness. Do not add irrelevant brands.
  4. JSON-LD schema -- place in page <head>, not only through client-side injection.
  5. Schema entity reinforcement -- mirror on-page entities inside schema fields (name, description, about, sameAs, knowsAbout, mentions, author, publisher).

Jesper Nissen's screenshot is a small field proof for why these rules stay bundled. The root post reports a homepage ranking jump after keyword placement, entity placement, entity-gap completion, and schema/entity reinforcement; the attached reply claims a CTR bump after applying the entity-gap tip and says schema is still underused. Treat this as anecdotal operational evidence for the existing checklist, not as a new ranking guarantee. Source: X/@JespernissenSEO and local image review, 2026-07-04

Master GEO Scorecard

Category Weight
AI citability and extractability 25%
Brand authority and off-site signals 20%
Content quality and E-E-A-T 20%
Technical foundations 15%
Structured data 10%
Platform optimization 10%

Bands: 90-100 elite, 75-89 strong, 60-74 mixed, 40-59 weak, 0-39 critical.

AI Citability Engineering

Models cite passages that are direct, self-contained, fact-rich, structurally legible, and make sense outside their original context. The 134-167 word passage target is a heuristic for high-value answer blocks.

5-part citability model: (1) Answer block quality -- question-based H2, direct answer in first 1-2 sentences, named entities, precise claims. (2) Self-containment -- names subject explicitly, no dangling pronouns, at least one concrete fact. (3) Structural readability -- one H1, clean H2-H3 nesting, short paragraphs, lists for steps, tables for comparisons. (4) Statistical density -- percentages, dollar values, dated claims, named studies, benchmarks. (5) Uniqueness -- original data, case studies, experiments, screenshots.

E-E-A-T

Experience: first-hand language, implementation notes, screenshots, case studies with numbers. Expertise: visible author, credentials, sourced claims, correct terminology. Authoritativeness: external citations, strong topic coverage across the site, media mentions. Trustworthiness: clear contact, privacy policy, accurate claims, visible dates, HTTPS.

Technical GEO Foundations

Crawlability: choose discovery, user-fetch, and training permissions separately. Google Search AI features use Googlebot and Search preview controls; Google-Extended affects specified Gemini training/grounding uses and has no effect on Google Search. ChatGPT search uses OAI-SearchBot while GPTBot is the separate possible-training control. Claude search uses Claude-SearchBot/Claude-User while ClaudeBot is the training crawler. Perplexity documents PerplexityBot/Perplexity-User as non-training search/fetch routes. Verify robots policy and CDN/WAF/server receipts before declaring a surface crawlable. Source: Google, OpenAI, Anthropic, and Perplexity first-party crawler documentation, 2026-08-11 review

Rendering: prefer SSR/SSG. Make schema server-rendered. Some AI crawlers are limited in JavaScript execution.

Core Web Vitals: LCP < 2.5s, INP < 200ms, CLS < 0.1.

Internal linking: high-value pages within ~3 clicks of homepage, descriptive anchor text.

Googlebot byte ceiling: Googlebot currently fetches up to 2MB for any individual URL, excluding PDFs, and that counter includes the HTTP header. If the HTML response exceeds that limit, Google processes the first 2MB as the fetched resource; referenced CSS/JS/image/video resources have their own fetch limits, and PDFs use a 64MB limit. Keep title, robots directives, canonicals, structured data, and primary content before bloated inline assets, menus, or hydration payloads. Source: Google Search Central, 2026-03-31; Source: local artifact review, 2026-07-03

Schema Design

Baseline graph: Organization/Person + WebSite + page-specific type + BreadcrumbList. Use sameAs to disambiguate across Wikipedia, Wikidata, LinkedIn, YouTube, GitHub, Crunchbase, etc. Prefer one @graph block with @id references. Use absolute URLs and ISO dates.

llms.txt

Optional machine-readable summary at domain root: what the site is, what sections matter, canonical pages, and core facts. Keep it aligned with the real site and test whether target agents consume it. It is not required for Google Search AI features, a universal ranking signal, or a replacement for useful content, ordinary crawlability, or semantic structure.

Competitive Comparison Capture

Capture branded commercial-intent queries directly. Publish exact-match vs, review, and alternatives pages. Make answer blocks so explicit that AI Overviews can lift them verbatim.

Per competitor: at minimum one standalone review page + one exact-match comparison page. Structure for answer extraction: question-led H2s, quotable verdict sentences, side-by-side tables, FAQ blocks.

Comparison page anatomy: exact-match title/slug/H1, first-paragraph verdict, Key Takeaways block, side-by-side table, pricing with dates, best-for guidance, honest tradeoffs, FAQ, clear CTA.

Firsthand testing required: sign up, use the products, take notes and screenshots. Fake reviews are obvious. Combine firsthand testing + external review language (G2, Capterra, Reddit) + your own positioning.

AIO conquest loop: pick 3-6 named competitors, ship one honest competitor-vs-product page and one standalone review per competitor, then inspect Google AI Overview / AI Mode / ChatGPT / Perplexity answers for those exact queries. The durable pattern from Kleo's bookmark is not "write attack pages"; it is firsthand use + external review language + extraction-friendly verdicts. Their screenshot shows Google returning a direct AI Overview answer for "kleo vs stanley" and citing Kleo's comparison/review pages. Source: X/@RobHoffman_, 2026-03-13

Local-service agent loop: for local businesses, route Claude/Cowork/browser agents through competitor page extraction, high-intent local keyword discovery, Google Business Profile completeness, service-page locality signals, review-response drafting, schema checks, and weekly rank/citation monitoring. The Claude Cowork bookmark is useful as a workflow sketch, but its attached images are only Claude and Google Business Profile logos; the durable asset is the checklist, not the visuals. Source: X/@bloggersarvesh, 2026-03-14

Programmatic SEO Domain Boundary

Programmatic SEO is a brand-authority decision, not only a template-generation tactic. Thin, short-term, or unproven pSEO experiments should not live on the primary brand domain: if they collapse, they can dilute topical trust, create crawl waste, and leave the brand with a footprint of disposable pages. Main-domain pSEO is reserved for durable pages with unique value, first-party or defensible data, useful internal links, clear indexation rules, and a long-term maintenance owner. Disposable experiments belong on a separate property, in noindex/staging, or outside the brand graph until quality and demand are proven.

The EXM row is useful as a blunt risk boundary rather than a precise traffic forecast: do not borrow the main site's authority for mass pages that are meant to be short-lived. Build authority and backlinks for the durable domain; keep churny pSEO tests isolated. Source: X/@EXM7777, 2026-03-12

Programmatic generation is not index admission

Jake Ward's resolved Byword article contributes a sound software boundary: versioned niche taxonomy -> strict JSON payload -> validation -> page-type renderer. It is better than freeform generated HTML because content, structure, and presentation can be tested and replaced independently. Its two utility questions are equally durable: would the page help without search engines, and would someone intentionally return to it? Source: X Article via raw/x-bookmarks/enriched/2031701565434732917.json, 2026-03-11

The reported outcome is not durable proof. The article self-reports 13,000+ pages, weekly clicks rising from 971 to 5,500 in 60 days, roughly half the pages indexed, and no negative Google signal. By 2026-08-11, Byword's own robots.txt said /resources/ had been retired during a Google “scaled content abuse” manual-action recovery and left crawlable so Googlebot could observe 410 Gone. Two matching resource examples returned 410; all 13 current sitemaps contained 863 URL entries. The current count does not recreate the earlier cohort, but it directly contradicts using the 60-day snapshot as a safety or permanence claim. Source: Byword robots.txt, Byword sitemap, and captured HTTP responses, 2026-08-11

The correct system separates three decisions:

generate candidate -> validate payload -> render and inspect
  -> stage/noindex cohort -> admit selected URLs to index
  -> monitor -> refresh, consolidate, noindex, or remove

Index admission needs page-level utility, source and rights provenance, factual checks, batch similarity and cannibalization checks, raw HTTP metadata and content, rendered behavior, accessibility/performance, an owner, an index budget, and a removal/repair path. Type-valid JSON proves shape, not value. Purpose-built UI can improve value, but an interactive checkbox or filter does not rescue generic or unsupported content.

Google's current policy is purpose- and value-based: many automatically or manually generated pages can be scaled content abuse when their primary purpose is ranking manipulation rather than helping users. For a manual action, fix all affected patterns, keep the relevant URLs reachable for review, document the quality issue and remediation, and request reconsideration only after the systemic cause is corrected. Source: Google scaled-content policy, generative-AI guidance, and Manual Actions report, reviewed 2026-08-11

Search Console Query Refresh Loop

Google Search Console is the cheapest refresh signal because it shows the queries a page already earns impressions for. Use it for pages that have traction but under-capture intent.

Workflow:

  1. Pick one page with meaningful clicks or impressions in the last 90 days.
  2. Export the page-filtered query CSV from Search Console.
  3. Ask Claude to group queries into satisfied intent, missing intent, title/H2 opportunities, FAQ opportunities, and internal-link opportunities.
  4. Update the page with answer blocks, missing subtopics, schema fields, and clearer headings.
  5. Re-measure clicks, impressions, CTR, and average position after the next crawl window.

The reviewed artifact shows the intended input shape: a Search Console performance view with hundreds of clicks, six-figure impressions, low CTR, and average position around page one. The durable lesson is not the screenshot metrics; it is the page-refresh loop from query evidence to editorial changes. Source: X/@ayushtweetshere, 2026-03-14

The reviewed @hridoyreh artifact sharpens the target selection: filter Search Console for high-impression queries where an existing page ranks around positions 8-20, then update that page with missing answer blocks, People Also Ask coverage, headings that match search intent, and a few internal links from stronger pages. This is a refresh loop for pages already near the prize, not a generic "publish more posts" tactic. Source: X/@hridoyreh and local image review, 2026-07-03

AI Search Measurement Loop

Treat AI-search visibility as a parallel measurement lane, not a replacement label for SEO. The primary John Mueller reply says businesses that earn referred traffic should inspect the full picture, actual usage metrics, and audience percentages before allocating time. It does not say Google ranking determines inclusion in ChatGPT, Claude, Gemini, or Perplexity, and it does not endorse the attached vendor packages. Source: John Mueller primary Reddit comment

Google's June 2026 Search Console release changes the operating loop. A subset of properties now has dedicated Search and Discover generative-AI reports with impressions, pages, countries, devices where supported, and time granularity. The data remains in the overall report. Use the dedicated report when available; retain overall Web/Performance and first-party conversion analytics as the baseline. Google explicitly warns that no third-party tool has access to its internal ranking or AI systems. Source: Google Search Console announcement; Source: Google generative-AI optimization guide

Evidence lane Record What it proves
Audience allocation Channel usage %, target segment, business value, time/cost allocation Whether the surface deserves work
Google Search AI Dedicated Search Console report when available; overall report fallback Google-reported impressions/pages/time trends, not another engine's visibility
Answer observation Exact query, product/model, date/time, locale, account state, mode, mention, cited URL, screenshot/export What one controlled run displayed
Stability Repeated trials over time with the same panel Volatility and directional change, not causality by itself
Business outcome First-party referral, engagement, signup, revenue, or other conversion event Whether observed visibility created value
Third-party tracker Disclosed query panel, collection method, timestamps, exports A vendor's reproducible observation—not internal rank or causal truth

The complete @alexgroberman replay reviewed all five SEO Stuff links and four images. The Search Console screenshot lacks property/date/filter context; six growth charts lack intervention and comparison data; one ChatGPT 5.2 answer is anecdotal; and the citation-column screenshot is a useful monitoring-interface pattern. Keep the offer as market signal and the interface as design signal. Hold the 60–90-day results, client counts, backlink causality, “safe to repeat,” and “single biggest factor” claims until independently reproducible. Source: X/@alexgroberman; local replay receipt: .brain/artifacts/x/2032084993368080515/seo-geo-audience/source-snapshots.md.

Platform-Specific Notes

Google Search AI features: ordinary Googlebot index/snippet eligibility, valuable non-commodity content, semantic structure, truthful structured data, and Search Console measurement. ChatGPT: verify OAI-SearchBot, keep important facts public and stable, and use a repeated prompt/referral panel. Perplexity: verify PerplexityBot and record the exact cited pages over repeated trials. Claude: verify Claude-SearchBot/Claude-User; decide ClaudeBot training access separately. Gemini: do not infer Google Search inclusion from Google-Extended or a Gemini answer.

Audit Workflow

  1. Define target (keyword, entity, page type, geography, platform mix)
  2. Fast SEO/GEO audit (title, H1, slug, intro, schema, canonical, dates, author, crawlability)
  3. Keyword and entity gap
  4. Score (citability, E-E-A-T, technical, schema, platform fit)
  5. Rewrite (title, H1, H2 outline, intro, answer blocks, schema)
  6. Strengthen graph (author pages, sameAs, external profiles, llms.txt)
  7. Deliver prioritized action plan

Agent Skills (Enforced)

Route SEO/GEO work through the skill registry rather than hard-coded home-directory projections:

  • ai-seo (skills/personal/ai-seo/SKILL.md) — AI-search visibility, source/crawler policy, repeated prompt panels, and measurement.
  • programmatic-seo (skills/personal/programmatic-seo/SKILL.md) — page-family strategy, typed generation, candidate/index separation, staged cohorts, and scaled-content recovery.
  • seo-audit — traditional crawlability, indexation, technical, on-page, content, and authority audit.
  • schema — accurate structured-data implementation when the visible content supports it.
  • og-metadata-audit — OpenGraph/Twitter card metadata and preview behavior.

Use wiki/meta/skill-registry.md and wiki/SKILL-RESOLVER.md for the current installed paths. This page owns the knowledge model; executable behavior belongs in the skill source and must be regenerated into runtime projections after edits.

OG / Social Preview Lessons

Learned from Loop (loooop.dev) cross-referenced against Dedalus monorepo (April 2026):

  • twitter: namespace is permanent - X/Twitter did not change the meta tag namespace. The crawler still reads twitter:card, twitter:image, etc.
  • Next.js middleware naming is strict - file must be middleware.ts, export must be middleware or default. A file named proxy.ts is silently ignored. This was the root cause of broken Twitter cards on Loop.
  • max-image-preview: "large" in root robots metadata tells Google (and helps Twitter) to render full-size image cards instead of thumbnails.
  • Canonical inheritance trap - In Next.js App Router, child pages that don't export alternates.canonical inherit the root layout's canonical (usually /). Every page needs its own.
  • htmlLimitedBots + middleware - Two complementary layers. Middleware intercepts social bots with custom HTML. htmlLimitedBots in next.config.ts serves lightweight HTML to SEO/AI crawlers. They overlap for social bots but htmlLimitedBots extends coverage.
  • DRY pattern - buildPageMetadata(input) returns complete { openGraph, twitter, alternates, title, description } from one function call. Eliminates 15-line boilerplate per page and ensures canonical consistency.

Twitterbot delivery invariant

When an owned project serves dynamic OG routes, its robots.txt or Next.js robots.ts should carry an explicit Twitterbot group for the shared aliases, even when User-agent: * already allows the whole site:

User-Agent: Twitterbot
Allow: /api/og
Allow: /api/og-home

Also allow the routes the project actually emits, such as /opengraph-image, /twitter-image, /og, or /og.png. This explicit group is defense in depth for projects whose wildcard rule blocks /api/; it is not proof that a social card will render and it does not invalidate a cached failed scrape. Diagnose the public delivery chain with a Twitterbot/1.0 user agent: the page must return 200, expose absolute og:image and twitter:image URLs, and the selected image URL must return 200 with an image content type. If those checks pass but the card is stale, deploy the metadata/image change and use a versioned image URL or new route to force a fresh scrape. In the July 27 Local Search audit, Twitterbot already received 200 under the wildcard allow rule while production still served the older metadata, proving that stale deployment or social cache—not robots alone—was the active failure mode. Source: User request; live Twitterbot delivery audit of local-search-xi.vercel.app, 2026-07-27

Tools

GEO-SEO-Claude (GitHub): open-source Claude Code GEO/AEO reference with /geo audit, citability, crawler, schema, brand, platform, report, and PDF routes. Keep it reference-only unless a future pass extracts a missing capability into Kevin's active ai-seo, seo-geo, seo-geo-optimization, or seo-audit skills. Source: GitHub/source review, 2026-07-03

Jam / SpreadJam: growth-channel agent platform that turns AI-search visibility and SEO/GEO audits into workflow actions. Use Kevin's installed SEO/GEO skills for direct audits; evaluate Jam when the task needs repeated growth-channel workflows, repo/site analysis, content-gap actions, warm leads, inboxes, or deliverability loops. Source: SpreadJam site/docs and local X video review, 2026-07-02

Cloudflare /crawl endpoint: single API call to crawl an entire site via Cloudflare Browser Rendering. Alternative to Firecrawl for automation use - compare latency, JavaScript rendering fidelity, and rate limits.

Signals (April 2026)

  • Googlebot 2MB fetch ceiling (Google Search Central, March 2026): Googlebot currently fetches up to 2MB for an individual URL, excluding PDFs, and the counter includes HTTP headers. The SEO implication is response-ordering, not "total page weight": keep metadata, canonical tags, schema, and primary content early in the HTML response, while external resources are fetched separately with their own limits. PDFs get 64MB. Source: X/@Charles_SEO local screenshot review, 2026-07-03
  • Claude Code SEO workflow (Cody Schneider, March 2026, 4K bookmarks): Set up a .env with Keywords Everywhere API key, DataForSEO API key, data warehouse for Google Search Console, and CMS access. Then use Claude Code for: keyword universe mapping (related keywords + SERP API for gap analysis), programmatic landing pages at scale (per-vertical with semantic terms + schema), link building (DataForSEO domain intersection for competitor backlink gaps), internal linking maps (topical relevance clusters from keyword data), and content refreshing. Source review resolved the self-thread to Graphed, a hosted marketing-agent service that positions this workflow as an autonomous growth operations system. Source: X/@codyschneider self-thread and Graphed site review, 2026-07-04
  • Free backlinks from high-authority domains: WordPress and Medium both offer free backlink surfaces - republish or cross-post carefully and set canonical links back to the source page when appropriate. The reviewed WordPress artifact shows a high-authority WordPress domain snapshot (DR 95, AR 31), so route this as authority-surface evidence, not a license to spray thin duplicate pages. Source: X/@kalashvasaniya and local image review, 2026-07-03
  • Google's John Mueller recommends tracking AI search traffic separately in analytics. No standard referrer header yet, but Perplexity and ChatGPT are increasingly identifiable.
  • Awesome Design MD (GitHub): reverse-engineered design systems of Apple, Spotify, Airbnb, and other companies as single markdown files. Useful as reference data for schema and entity modeling on design-related pages.

Taxonomy

12 operational categories: Foundations, Research & Intent, Site Architecture, Technical Crawl & Indexing, On-Page Entities & Schema, Content Systems & Page Types, GEO/AEO/AI Platforms, Authority Links & Brand, Distribution & Feedback, Measurement & Experiments, International, Growth Systems. Cross-cutting: UX, conversion, accessibility, legal, performance, editorial quality.

Hridoy Rehman's folder-style SEO bookmark is useful as a curriculum-shaped view of the same taxonomy: foundations, research, SERP analysis, search intent, competitor research, and keyword gaps belong before tactical page edits. Because the enriched record has no resolved source beyond the tweet and is visibly truncated, it is supporting structure for this page rather than a separate SEO skill or a claim about a complete course outline. Source: X/@hridoyreh, 2026-03-15


Timeline