elseweb

Crawler policy

elseweb fetches approved homepages, declared or moderator-approved feeds, and public article URLs supplied by those feeds. Fetching runs in the worker process.

Feed language declarations and publisher tags help classify articles. Language detection runs locally. When enabled, unresolved article categories use the operator’s configured AI service or local model with only the retained public feed title, excerpt and tags. Moderators can correct article languages and categories.

Article text supplied by feeds is indexed first. When a feed supplies only a summary, the worker can fetch the original article page. It respects robots.txt and indexing restrictions it encounters, and does not execute scripts or crawl links from articles. Requests share the existing destination, response-size, concurrency, and daily-budget controls.

We retain up to 100,000 characters of plain text per article for search, with short excerpts linking to the original publisher. Raw page HTML is not retained or rendered. Page text is normally rechecked weekly; failures retry after a day. Coverage depends on publisher access, feed content, and the crawl budget. Removal clears retained article bodies.