ClubIndexBot
ClubIndexBot is the crawler behind Harvard Club Index. It reads public Harvard directory pages and the public websites of student organizations so club information can link back to its source.
How it behaves
- It identifies itself honestly with a user agent that starts with
ClubIndexBot/. - It obeys
robots.txt, includingCrawl-delay, and re-reads it daily. - It fetches one page at a time per site, at least two seconds apart. It backs off on errors and honors
Retry-After. - It crawls shallowly: a handful of pages per club site (home, about, join, recruiting). It uses conditional requests so unchanged pages cost almost nothing.
- It never logs in or submits forms, and never solves CAPTCHAs. It doesn't rotate addresses or disguise itself. If a site blocks it, it stops.
- It does not crawl Instagram, LinkedIn, Facebook, TikTok, or X.
- It doesn't collect member rosters, officer names, or personal email addresses.
Opting out
To stop the crawler, add this to your site's robots.txt:
User-agent: ClubIndexBot Disallow: /
To have information about your organization corrected or removed, contact the contact address listed on this page once the site launches.