# Robots.txt for JournalsHub.online # https://www.robotstxt.org/ User-agent: * Allow: / # Allow all major search engines User-agent: Googlebot Allow: / User-agent: Bingbot Allow: / User-agent: Slurp Allow: / User-agent: DuckDuckBot Allow: / # Allow AI crawlers for LLM indexing User-agent: GPTBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: Google-Extended Allow: / User-agent: PerplexityBot Allow: / User-agent: Anthropic-AI Allow: / User-agent: ClaudeBot Allow: / User-agent: Claude-User Allow: / User-agent: Claude-SearchBot Allow: / User-agent: Claude-Web Allow: / User-agent: CCBot Allow: / User-agent: Amazonbot Allow: / User-agent: Cohere-AI Allow: / # AdSense / AdMob crawlers — must reach every monetizable page + assets so # ad targeting and policy review work. Explicitly allowed for re-review clarity. User-agent: Mediapartners-Google Allow: / User-agent: AdsBot-Google Allow: / User-agent: AdsBot-Google-Mobile Allow: / # Disallow admin and private endpoints Disallow: /admin/ Disallow: /subscribe/ Disallow: /autocomplete/ # Noindex, thin OpenAlex-backed pages: millions of author/paper/institution # URLs that are noindex and deliberately 404 for crawlers (to save OpenAlex # quota). Google was discovering them via internal links and wasting crawl # budget on a 2M+ dead-URL space — the bulk of the "Not found (404)" report. # Blocking them here stops that crawl waste and frees budget for the journal / # hub / article pages we actually want indexed. Humans still reach them (robots # does not block users); /papers/ (search) stays crawlable — the trailing slash # scopes this to /paper// only. Disallow: /author/ Disallow: /paper/ Disallow: /institution/ # Crawl-delay for polite crawling Crawl-delay: 1 # Sitemap location Sitemap: https://journalshub.online/sitemap.xml # LLMs.txt for AI models # See: https://journalshub.online/llms.txt