Developer news without the noise — My Dev News in Telegram.

Open botLearn more
Stuzhuk Lab

Stuzhuk Lab — Chemistry of Code

Chemistry of Code

  • Home
  • Resume
  • Portfolio
  • Services
  • Blog
  • Contact
  • Sign in
  • 🇬🇧
    🇷🇺🇺🇦
← All posts

Tag

evaluation

All blog posts with this tag.

  • 2 Sept 2026

    Dataset engineering for model training: where to start

    A pillar guide to data for training and evaluating AI: golden records, golden editors, train/eval splits, leakage, manifests, and parameters by task and domain.

    datasetsevaluationfine tuningmachine learningml dataset engineering
  • 21 Aug 2026

    Language Model Quality Testing: What the Tests Are and What They Actually Measure

    LLM quality checks: public task sets, arenas, model judges, safety, product regression, and online signals — what each test measures and what it cannot see.

    AIarchitectureevaluationllmqualitytesting
  • 1 Aug 2026

    AI Evaluation Harness in 2026: Golden Sets, Regression Gates, and Release Criteria That Survive Production

    Build an AI evaluation harness that connects golden datasets, retrieval and generation metrics, CI regression gates, shadow traffic, release controls, observability, and LLM gateway FinOps.

    AIarchitectureevaluationobservabilityragtesting
  • 25 Jul 2026

    Evaluating Enterprise AI: Metrics You Can Trust

    A practical framework for measuring enterprise AI across retrieval, answers, security, and business outcomes.

    ai enterprise architectenterprise aievaluationragsecurity