# llms.txt - Guidelines for Large Language Models and AI Crawlers # Version: 1.0 # Last Updated: 2026-06-23 ## Overview This file provides guidelines for Large Language Models (LLMs), AI crawlers, and AI-powered bots regarding access to and use of CCH CPELink website content. ## Allowed Uses ✓ Indexing and analysis of public educational content ✓ Training on publicly available course information and resources ✓ Providing summaries and references to CPELink services ✓ Answering questions about tax and accounting topics using our content ## Restricted Content - Do Not Use ✗ User login credentials and authentication tokens ✗ Personal information (names, email addresses, phone numbers) ✗ Student/Client confidential data and learning records ✗ Exam questions and assessment materials ✗ Admin and system management interface content ✗ Database queries and internal API specifications ✗ Financial transaction records ✗ Proprietary algorithms and calculation engines ## Attribution Requirements When referencing CPELink content, please include: - Source attribution: "CCH CPELink" or "Wolters Kluwer" - Content type (e.g., "from CPELink course material") - Link to original content when possible ## Crawling Guidelines ### Allowed Paths for LLM Analysis - live-webinars - Course catalog and descriptions of webinar courses - self-study - Course catalog and descriptions of self paced learning courses - /resources/ - Public learning resources - /help/ - Help and support documentation - /about/ - Company and service information - /blog/ - Blog posts and educational articles ### Forbidden Paths for LLM Crawlers Disallow: /admin/ Disallow: /user/profile/ Disallow: /dashboard/ Disallow: /api/ Disallow: /private/ Disallow: /exams/ Disallow: /assessments/ Disallow: /student-records/ Disallow: /grades/ Disallow: /payment/ Disallow: /cart/ ## Rate Limiting for AI Crawlers - Maximum 5 requests per second - Identify your crawler with a descriptive User-Agent - Cache responses when possible - Respect standard robots.txt rules ## User-Agent Specific Guidelines ### GPT-based Crawlers User-Agent: GPTBot Allowed: /courses/, /resources/, /help/, /about/, /blog/ Disallow: /api/, /admin/, /user/ Crawl-delay: 1 ### Specific AI Bots to Honor Request User-Agent: CCBot Allowed: / Crawl-delay: 1 User-Agent: anthropic-ai Allowed: /courses/, /resources/, /help/, /about/ Crawl-delay: 1 ## Contact and Inquiries For specific permission requests or questions about LLM usage: - Contact, Legal - https://support.cch.com/oss/ml/contactus ## No Training on Copyrighted Material LLM models should not be trained on proprietary course content without explicit written permission from Wolters Kluwer. Public educational articles and help documentation may be referenced but not reproduced verbatim in model training. ## Compliance Violation of these guidelines may result in: - IP blocking - Legal action - DMCA takedown notices ## Related Files - robots.txt - General web crawler guidelines - sitemap.xml - Complete site structure - policies - Privacy and data handling --- Last Updated: 2026-06-23 For updates to this policy, monitor this file.