Customer Support English Crawling

The Customer Support English Crawling is a 2-million-token corpus of English built from the web by targeting specific in-domain urls that belong to the customer support domain such as FAQ and help websites, as well as community sites and forums. It consists of 2,025,813 tokens, 169,344 sentences and 5,636 documents.
Documents are separated by single new lines.
The corpus has been developed in the framework of the CEF project MT4ALL (
We license the actual packaging of this data under a CC0 1.0 Universal License.