# ASP-RCM Solutions - robots.txt # https://asprcmsolutions.com/robots.txt # Updated 2026-08-12 (AI crawlers un-blocked; was 2026-06-25 crawl-efficiency pass) # # Goal: let Googlebot reach EVERY real page (incl. all ~9,100 city pages # and the full sitemap) while keeping it out of parameter/junk URL spaces # and search/filter traps that waste crawl budget. User-agent: * Allow: / # --- Asset dirs and internal files (not content; no need to crawl) --- Disallow: /404.html # CSS, JS and images must stay crawlable: Google renders the page with them, # the nav/footer come from /js/partials.js, and /img/ feeds the image sitemap. Allow: /css/ Allow: /js/ Allow: /img/ Disallow: /_includes/ Disallow: /_proof/ Disallow: /_headers Disallow: /_redirects Disallow: /vercel.json Disallow: /netlify.toml Disallow: /.htaccess Disallow: /SEO-PLAYBOOK.md Disallow: /SEO-FIXES-2026-06-20.md Disallow: /LAUNCH-CHECKLIST.md Disallow: /README.md Disallow: /search-index.json # --- Search / filter traps ------------------------------------------- # The on-site search uses ?q= and ?id= which spawn unlimited URL # variations. Block the crawlable forms so Googlebot stops expanding # them. The /search/ page itself stays out of the index. Disallow: /search Disallow: /search/ Disallow: /*?q= Disallow: /*?id= Disallow: /*&q= Disallow: /*&id= # --- Generic parameter / junk URL spaces ----------------------------- # Tracking params, pagination noise, sort/filter combinations, and # common CMS-bot probe paths. These never point to unique content, so # blocking them concentrates crawl budget on real city pages. Disallow: /*?utm_ Disallow: /*?fbclid= Disallow: /*?gclid= Disallow: /*?ref= Disallow: /*?sort= Disallow: /*?filter= Disallow: /*?page= Disallow: /*?s= Disallow: /*?replytocom= Disallow: /*?*sessionid= Disallow: /wp-admin/ Disallow: /wp-login.php Disallow: /xmlrpc.php Disallow: /feed/ Disallow: /*/feed/ Disallow: /cgi-bin/ # NOTE: City pages are clean, parameter-free URLs of the form # /specialties/-billing-services// # /specialties/-billing-services/states// # None of the Disallow rules above match those, so every city page # remains fully crawlable. # --- AI crawlers: WELCOME (policy set 2026-08-12) -------------------- # Business decision: AI crawlers (OpenAI/GPTBot, Anthropic/ClaudeBot, # Google-Extended, PerplexityBot, CCBot, OAI-SearchBot, ChatGPT-User, # Perplexity-User, Applebot-Extended, etc.) are ALLOWED for training AND # live retrieval/citation, to support AI-search visibility in ChatGPT, # Claude, Perplexity, and Google AI Overviews. See /llms.txt. # They are intentionally NOT given their own groups: per RFC 9309 they # fall through to the "*" group above and inherit the SAME full trap/asset # blocklist Googlebot uses - full access to real pages, zero crawl-budget # waste on traps. (This replaced a prior Disallow-all block that # contradicted /llms.txt.) # # Bytespider (ByteDance) is the one aggressive scraper we still fence off, # since it hammers origin without driving AI-search citation: User-agent: Bytespider Disallow: / # --- Sitemaps (use the canonical non-www host, matches sitemap ) - Sitemap: https://asprcmsolutions.com/sitemap.xml Sitemap: https://asprcmsolutions.com/sitemap-index.xml