
SFT datasets: designing a supervised fine-tuning corpus
A pillar on instruction–response corpora for SFT: pair contracts, coverage, refusals, synthetic vs expert sources, pre-train quality gates, and leakage with evaluation sets.
Developer news without the noise — My Dev News in Telegram.

Tag
All blog posts with this tag.

A pillar on instruction–response corpora for SFT: pair contracts, coverage, refusals, synthetic vs expert sources, pre-train quality gates, and leakage with evaluation sets.

A pillar guide to data for training and evaluating AI: golden records, golden editors, train/eval splits, leakage, manifests, and parameters by task and domain.

A practical map of LLM adaptation: what changes model weights, what plugs in knowledge and tools, and how to choose an approach without “fine-tuning on PDFs.”

How to separate enterprise knowledge from model behavior and decide when fine-tuning is worth the operational cost.