Hugging Face ML-intern Builds Six Custom AI Models for Just $103
Hugging Face's ML-intern built six specialized AI models for about $103, automating data, training, evaluation and publishing under strict cost controls.
Summary
On October 8, 2026, Hugging Face’s ML-intern turned detailed HuggingChat briefs into six public Hub models, planning jobs, requesting spending approval, smoke-testing, training, evaluating and publishing model cards on Hugging Face hardware. Budgets start at $0 and remain capped. Seven 450 to 2,000-word prompts at yvrjsharma/ml-intern-prompts specify datasets, base models, scripts, verified facts, pretraining baselines and weight-change checks. Reported CPU and GPU charges totaled about $103.
Citrus Doctor used 3,017 images across 21 conditions; Qwen3.5-2B rose from 14.9% to 52.8% on 335 photos after two A10G epochs costing $1.90. The $7.60 Huggy LoRA trained FLUX.2 klein base 4B on 84 drawings; checkpoint 200 worked, while style leakage began after step 500, and distilled generation takes four steps. Viewpoint Orbit rendered 1,030 objects at 24 angles, creating 24,722 images; 461 training objects and 40 tests produced 1,844 pairs across 23 instructions before 2,000 A100 steps and 48 jobs cost $16.
Doodle-in built 6,042 pairs and 160 tests using Open Images and LaMa while recording licenses; checkpoint 500 delivered 67.5% placement in 4.7 seconds, with unseen classes scoring 65.0% versus 64.2%, across 59 jobs costing $24.30. Pocket Rewriter replaced a 9B, 20 GB, thousands-token Qwen-Image 2.1 rewriter: 8,797 labels yielded 1,840 examples and 0.8B and 2B students for $16.05; the 812 MB CPU GGUF returns valid output 99.7% of the time using one-quarter the tokens. Agate distilled the 260M Preview 002 from 50 guided steps, or 100 passes, to four; 155,000 cached images plus 24,000 teacher pairs lifted GenEval from 0.509 to 0.536 versus 0.563, using one-quarter compute. Two runs cost $37 and produced browser-ready ONNX.
Positives
- Qwen3.5-2B fine-tuning lifted citrus diagnosis accuracy from 14.9% to 52.8% on 335 photos for about $1.90.
- Pocket Rewriter’s 812 MB CPU model achieved 99.7% valid output while using roughly one-quarter of the 9B teacher’s tokens.
- Doodle-in generalized to 23 unseen object classes, scoring 65.0% placement versus 64.2% for the remaining classes.
- Agate reached 0.536 GenEval in four steps, approaching the 50-step teacher’s 0.563 with one-quarter of the compute.
- ML-intern enforced approved budgets while automating dataset creation, smoke tests, training, evaluation and public Hub publishing.
Risks & concerns
- FLUX.2 klein’s Huggy style began leaking into unrelated prompts after checkpoint 500, indicating overtraining.
- Viewpoint Orbit needed 48 jobs because missing packages and incorrect paths caused failures and resubmissions.
- The original Qwen-Image 2.1 rewriter requires about 20 GB of memory and thousands of reasoning tokens before producing one paragraph.
- Agate’s four-step GenEval score of 0.536 remained below the 50-step teacher’s 0.563 despite two training runs.
- Doodle-in placed 67.5% of requested objects correctly, leaving nearly one-third undetected at the drawn location.