AI Dataset Engineering

Image annotation

Image annotation labels what a vision model must learn: classes, boxes, polygons, or tags on product photos, defects,…

Discuss this directionContact form

Tell me the goal, stack constraints, and timeline — I reply on Telegram.

Image annotation labels what a vision model must learn: classes, boxes, polygons, or tags on product photos, defects, shelves, or scans you own. I write the guideline before a large batch starts, so two people do not invent two meanings for the same class. Image annotation is dataset work. Reading fields off invoices in production is document recognition, not this page.

Volume is a separate line in the estimate. A pilot labels a sample you can review; a full batch follows only after the guideline survives that review. I do not promise a crowd of annotators inside the engineering hours. Exports land next to the images with stable file ids. Near-duplicate photos are handled in deduplication so the same object is not both train and validation.

Acceptance is the class list, the guideline, a double-checked sample, and files a training job can join back to the images. A guideline plus a review sample is often 2–4 weeks. The full batch depends on how many images you actually need.

Acceptance criteria

Done when

  • Schema and field meanings are written down
  • Train / validation / test split is reproducible and checked for leakage
  • Quality report lists counts, removed duplicates, and known gaps

Deliverables

  • Dataset files in the agreed format
  • Reproducible preparation script
  • Quality report

Out of scope

  • Training the model and production deployment
  • Legal opinion on personal data and third-party licenses
  • Annotator volume beyond the agreed sample unless it is in the quote

The final acceptance checklist is confirmed in the brief or contract; the list above is a scope alignment guide.

Ballpark estimate

Scope size
Extras

FAQ

Tap a question to expand the answer.

Do you label in our tool?

Yes, if the tool already exists and can export. I do not force a new platform for a few hundred images.

Is OCR included?

Field OCR for business documents is a different service. Here the label is what you define for training, not a form parser.

Discuss this directionContact form

Tell me the goal, stack constraints, and timeline — I reply on Telegram.

View full service page