Data collection and consolidation turns scattered company files into one corpus with a known origin. I pull the exports you already own — CRM, tickets, chats, spreadsheets, object storage, scans — and record where each row came from, when it was taken, and who may use it. Collection and consolidation is the first step of AI Dataset Engineering: without a source map, later cleaning and deduplication cannot be repeated.
I do not scrape systems you do not own, and I do not buy a third-party dataset to fill gaps unless that purchase is a line in the brief. Tokens stay in environment secrets. If a source has no export, I say so instead of automating a browser and calling it a pipeline. Joins use a stable key you can explain, not a lucky column name.
Acceptance is a source inventory, a sample joined on that key, and a script that rebuilds the consolidated table from the same exports. A single-source pull is often 1–3 weeks after access exists. Cleaning starts only after this table is real — see cleaning and normalization.
