AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: An AI Model Of Your Own: From Need To Build on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A Hugging Face contributor reports using its ML-intern agent to build and publish seven custom models over several days. Examples include a CPU-friendly prompt rewriter and a citrus disease classifier; performance, costs and workflow details are the contributor’s own account, not an independent evaluation.

A Hugging Face contributor says they used the platform’s ML-intern agent to build and publish seven custom models over several days, as described in the original analysis, including a smaller prompt rewriter and a citrus-image classifier. The account describes a workflow in which the agent proposed training plans, requested spending approval and ran evaluations, but its performance figures and costs are self-reported and have not been independently verified in the supplied material.

The first project addressed a practical limitation: the contributor wanted a smaller version of the prompt rewriter included with Qwen-Image 2.1. They said the original model has 9 billion parameters, needs about 20 GB of memory and can generate thousands of tokens before producing a paragraph. After finding compressed versions of the large model but no smaller alternative, they used a 9B model to label 8,797 example requests and trained a 0.8B model, a process related to model distillation. The contributor reported valid output 99.7% of the time and token use at about one-quarter of its teacher’s, with compute costing about $16.

Another project fine-tuned Qwen3.5-2B to identify citrus pests, diseases and nutritional deficiencies in images. The contributor says the dataset had 3,017 annotated images across 21 categories, illustrating the custom-model training workflow also discussed in coverage of IBM’s Granite model. On 335 test photos, they reported that accuracy on the correct problem rose from 14.9% for the untuned model to 52.8% after two training epochs on one A10G GPU. The reported compute bill was about $1.90.

The account also describes a character-generation LoRA trained on 84 captioned drawings and a camera-angle LoRA for Qwen-Image 2.1. For the latter, the agent generated 24,722 transparent images of household objects across 24 angles; selected objects were used for training and others held back for testing. The contributor said training took about 90 minutes on one A100 and the project cost about $16 in compute, including failed jobs that had to be resubmitted. The source gives detailed descriptions for only some of the seven models.

At a glance
reportWhen: Reported in an account published by Tho…
The developmentA Hugging Face contributor has described using the ML-intern agent to plan, train, evaluate and publish seven custom models over several days.
At a glance
reportWhen: Reported last week; the projects were b…
The developmentA Hugging Face contributor says an AI agent called ML-intern helped plan, train, evaluate and publish seven custom models on the Hub over several days.

What Lower-Cost Model Building Could Change

The report offers an example of how an agent could reduce the coordination work involved in building a specialized model. Instead of manually arranging each stage, a user with a specific task can ask an agent to plan work, run a baseline, test a training setup and prepare an evaluation. If the reported workflow performs reliably for other users, it could make small-scale customization more accessible to developers and organizations that do not have a dedicated machine-learning team.

The examples also show why a model’s score needs a comparison. The citrus result is presented against the base model on the same test set, giving readers a stated reference point. But a reported improvement on one set of 335 photos does not establish how the classifier would perform on different crops, lighting, devices or regions. Likewise, a low compute bill is not a full measure of project cost: the figures do not account for all data preparation, user review or other expenses.

Customization can also create unwanted effects. The contributor said later checkpoints for the character LoRA began influencing prompts unrelated to the intended character. That observation suggests a practical risk of training too long or too narrowly: a model may apply a learned trait more broadly than intended. It is a report from one project, not a general measurement of how often that occurs.

Amazon

AI model training hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Contributor Used ML-Intern

According to the account, each project began as a message in HuggingChat with ML-intern enabled. The agent proposed a plan and, before paid work, asked for spending approval. When a prompt lacked a budget, it offered options and asked the user to choose. The contributor says the agent then handled steps including training, evaluation and publication using Hugging Face hardware.

The prompts became more detailed as the contributor gained experience: they grew from about 450 words for the first project to nearly 2,000 by the sixth. They included the dataset, base model and training script, along with requests for verified facts, a baseline, a smoke test and a spending cap. The contributor says all seven prompts are available in a public GitHub repository and that model cards and evaluations were published on the Hugging Face Hub.

The account is a description of one contributor’s use of the tool, not an independent assessment of ML-intern. The contributor’s suggested safeguards included measuring the base model before training, running a small test and checking that saved weights changed. The source does not establish that those steps were used identically across all seven projects.

“Also report the base model’s zero-shot score on the same metric before training so we can see the gain.”

— The Hugging Face contributor, in a prompt described in the account

Amazon

GPU for machine learning

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Report Does Not Establish

The reported results have not been independently replicated in the supplied account, and detailed evaluation methods are not provided for every model. It is unclear how the contributor measured the prompt rewriter’s 99.7% valid-output rate, whether test images were independently reviewed, or how results would change on data outside the stated test sets. The account also does not provide full descriptions of all seven projects.

The cost figures refer to reported compute spending, not a complete accounting of development time, dataset preparation or review. It is also not clear whether other users would see similar costs or performance, or how the agent handled data quality across projects. The contributor’s account does not amount to a guarantee of results for other tasks, budgets or hardware.

Amazon

AI model distillation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Where Readers Can Check the Work

The contributor says the seven models and their evaluations are available on the Hugging Face Hub, with the project prompts posted in a public GitHub repository. Those materials may let readers inspect the published model cards and compare the reported evaluation information with the underlying projects, where datasets and other details are provided.

Further evidence would come from independent tests across more users, datasets and tasks, with consistent reporting of baselines, evaluation methods and total costs. Until then, the account provides a concrete example of one person’s workflow with ML-intern, while the broader reliability and typical cost of agent-assisted model building remain open questions.

Amazon

custom AI model development kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is ML-intern in this report?

It is the Hugging Face agent the contributor says they used through HuggingChat to plan model projects, request spending approval, run training and evaluations, and publish results.

What models did the contributor describe?

The account gives details about a 0.8B prompt rewriter, a citrus-image classifier, a character-generation LoRA and a camera-angle LoRA. It says seven models were published, but the supplied account does not describe all seven in equal detail.

How much did the projects cost?

The contributor reported about $1.90 in compute for the citrus classifier and about $16 for both the prompt rewriter and camera-angle LoRA projects. These are self-reported compute costs, not complete project budgets.

Are the reported performance results independently verified?

Not in the supplied account. The figures come from the contributor, and the source does not provide independent replication or full evaluation details for every model.

Where can readers inspect the projects?

The contributor says the models and evaluations are on the Hugging Face Hub and that all seven project prompts are in a public GitHub repository.

Primary source: Hugging Face · via ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Solving The Jane Street Reverse Engineering Challenge

Jane Street reportedly solved its complex reverse engineering challenge, marking a significant milestone in algorithmic problem-solving and cryptography.

Anthropic Says Its Biology Lab Has Already Found Something Big – TechCrunch

Anthropic says its biology lab has found something significant, but the available account gives no details or evidence to assess the claim.

Atari Surges In Global Coverage

Search interest and media mentions of Atari spike significantly, with 11 mentions in recent coverage, indicating rising global attention. The cause remains unconfirmed.

How Anthropic’s AI Watermark Outpaces Competitors — For The Time Being

Anthropic currently outpaces competitors by watermarking Claude’s responses, but the technology’s durability and adoption are still uncertain amid regulatory and technical challenges.