# # robots.txt # # This file is to prevent the crawling and indexing of certain parts # of your site by web crawlers and spiders run by sites like Yahoo! # and Google. By telling these "robots" where not to go on your site, # you save bandwidth and server resources. # # This file will be ignored unless it is at the root of your host: # Used: http://example.com/robots.txt # Ignored: http://example.com/site/robots.txt # # For more information about the robots.txt standard, see: # http://www.robotstxt.org/robotstxt.html User-agent: * # CSS, JS, Images Allow: /core/*.css$ Allow: /core/*.css? Allow: /core/*.js$ Allow: /core/*.js? Allow: /core/*.avif Allow: /core/*.gif Allow: /core/*.jpg Allow: /core/*.jpeg Allow: /core/*.png Allow: /core/*.svg Allow: /core/*.webp Allow: /profiles/*.css$ Allow: /profiles/*.css? Allow: /profiles/*.js$ Allow: /profiles/*.js? Allow: /profiles/*.avif Allow: /profiles/*.gif Allow: /profiles/*.jpg Allow: /profiles/*.jpeg Allow: /profiles/*.png Allow: /profiles/*.svg Allow: /profiles/*.webp # Directories Disallow: /core/ Disallow: /profiles/ # Files Disallow: /README.md Disallow: /composer/Metapackage/README.txt Disallow: /composer/Plugin/ProjectMessage/README.md Disallow: /composer/Plugin/Scaffold/README.md Disallow: /composer/Plugin/VendorHardening/README.txt Disallow: /composer/Template/README.txt Disallow: /modules/README.txt Disallow: /sites/README.txt Disallow: /themes/README.txt # Paths (clean URLs) Disallow: /admin/ Disallow: /comment/reply/ Disallow: /filter/tips Disallow: /node/add/ Disallow: /search/ Disallow: /search? Disallow: /user/register Disallow: /user/password Disallow: /user/login Disallow: /user/logout Disallow: /media/oembed Disallow: /*/media/oembed # Paths (no clean URLs) Disallow: /index.php/admin/ Disallow: /index.php/comment/reply/ Disallow: /index.php/filter/tips Disallow: /index.php/node/add/ Disallow: /index.php/search/ Disallow: /index.php/search? Disallow: /index.php/user/password Disallow: /index.php/user/register Disallow: /index.php/user/login Disallow: /index.php/user/logout Disallow: /index.php/media/oembed Disallow: /index.php/*/media/oembed # D-SEO-5 (owner ruling 2026-08-25): contain the /business facet URL space. # # Modelled on doctornet's DG-196 block, but every rule below is re-derived from # diakopesnet's OWN configuration and measurements -- per D23, never inherit # another portal's markers or numbers. What that re-derivation changed: # # * doctornet blocks `?*dist=`. diakopesnet has NO proximity/radius filter at # all (views.view.search_business, checked 2026-08-25) -- that rule is NOT # copied here. # * diakopesnet exposes SEVEN facet markers, not the five its own route table # assumed: c (category) and r (region) are multi-select OR links; e, f, p, v, # w are boolean AND checkboxes -- Has email / Facebook / photos / video / # website (facets.facet.*, field_identifier field_*_exists). # * doctornet keeps two facet segments crawlable because its sitemap advertises # ~21,500 combination URLs. diakopesnet's advertises ZERO facet URLs and # 63,391 node detail URLs instead. Two segments stay crawlable here for a # different reason: /categories and /regions are noindex,FOLLOW (D-SEO-3) and # reach the detail pages THROUGH these listings, so this is the gateway. # # Every URL below is already "follow, noindex" (34 probes, p3.3-seo-parity-d11), # but noindex has never stopped a crawl -- a crawler must fetch and render a page # to read that tag. robots.txt is what stops the fetch. The known trade-off: a # disallowed URL's noindex is never read, so an externally linked one can still # be indexed URL-only. Accepted, as on doctornet. # # `page=` is deliberately NOT blocked -- the master plan names # "business/c/, incl. ?page=N" explicitly. # Three or more stacked facet pairs (6+ path segments after /business). Disallow: /*/business/*/*/*/*/*/* # Same-axis stacking: c/a/c/b and c/b/c/a are the same OR-union in two orders. Disallow: /*/business/c/*/c/ Disallow: /*/business/r/*/r/ # The five boolean facets -- pure filters over an existing set, no unique content. Disallow: /*/business/e/ Disallow: /*/business/f/ Disallow: /*/business/p/ Disallow: /*/business/v/ Disallow: /*/business/w/ Disallow: /*/business/*/e/ Disallow: /*/business/*/f/ Disallow: /*/business/*/p/ Disallow: /*/business/*/v/ Disallow: /*/business/*/w/ # Exposed sort permutations (views.view.search_business exposes sort_by, # sort_by_1 and sort_by_2, all in url.query_args cache contexts). Disallow: /*?*sort_by= # /business-map is noindex and renders 2,500 entries in 796 KB per request # (measured 2026-08-25) -- the most expensive page on the site, zero crawl value. Disallow: /business-map Disallow: /*/business-map # D-SEO-1 (owner ruling 2026-08-25): robots.txt must point at the sitemap. # # Neither www.diakopesnet.gr nor this repo's scaffold copy carried a Sitemap: # directive -- measured 2026-08-25, both empty. The sibling portal already ships # exactly these two lines in production (www.doctornet.gr/robots.txt), so this # is the proven shape, not a new invention. # # Absolute production URLs are correct here even though this file is served # locally and on the interim: `web/robots.txt` is a Deployer SHARED file # (apps/deploy/diakopesnet.yaml), so the deployed copy lives in shared/ and is # seeded per app. The interim's shared copy must NOT advertise these -- it gets # its own crawl-blocking copy in Phase 4. Sitemap: https://www.diakopesnet.gr/el/sitemap.xml Sitemap: https://www.diakopesnet.gr/en/sitemap.xml