ML Intern - Hugging Face Research Agent

Keep and recommend as a high-signal Hugging Face research capability: it can connect papers, citations, datasets, current code, sandboxes, Jobs, metrics, and ablations. Do not globally install or run it on sensitive work until its trace-upload and headless-authority defaults are explicitly contained.

What it is

ML Intern is an Apache-2.0 Python agent and web app from Hugging Face for researching, implementing, running, evaluating, and shipping ML work. Its useful contribution is not “paper chat.” It gives the agent separate research context, paper/citation and dataset tools, current GitHub/docs lookup, local or HF Space execution, HF Jobs, monitoring, approvals, persistence, and an autonomous research loop. Source: huggingface/ml-intern source at 550a209701701e6a9ac7cac70b8dbd508822d467, reviewed 2026-08-10

Its research prompt starts from anchor papers, follows downstream citations, reads methodology/experiments/results, ties each result to the exact recipe, validates dataset format, then finds current implementation code. The execution prompt adds a strong practical sequence: inspect data, preserve the requested scope, test the exact source in a representative sandbox, run one pilot before a batch, monitor metrics, diagnose failures, ablate, evaluate, and persist the result. That sequence now lives independently in Research Experiment Workflow; the workflow does not depend on this vendor or runtime. Source: agent/tools/research_tool.py and agent/prompts/system_prompt_v3.yaml, reviewed 2026-08-10

Current source snapshot

Surface Reviewed state
Repository huggingface/ml-intern, untagged main at 550a209701701e6a9ac7cac70b8dbd508822d467; 490 commits; 10,725 stars and 1,170 forks in the 2026-08-10 GitHub API snapshot
Package 0.1.0, Python >=3.11, entrypoint ml-intern = "agent.main:cli"
License Apache-2.0
Runtime LiteLLM/HF Inference Providers or OpenAI-compatible local endpoints; built-in HF research/data/repo/jobs/sandbox tools; FastMCP; CLI and web app
Tests uv run --extra dev pytest tests/unit -q: 508 passed, 174 warnings
Releases No public tags or GitHub releases found
Preserved evidence Full source, GitHub snapshot, original 16.58-second X video, contact sheet, probe, hashes, and audit under .brain/artifacts/x/2046543093856412100/ml-intern/

The repository HEAD is unchanged from the July review, but the current audit is materially deeper: source behavior, defaults, authority, trace payloads, media, and all unit tests were inspected rather than inferring the product from its README. Source: local frozen source and capture manifest, 2026-08-10

Capability map

Capability Implementation reality
Literature HF Papers, arXiv/ar5iv, Semantic Scholar search/citations/recommendations, methodology-section reading, linked-resource discovery
Data Hub dataset inspection, schema/split/sample checks, private/gated access with HF token
Code and docs HF docs/API lookup, GitHub repo/example/file tools, web search, MCP extension
Execution Local shell/filesystem by default; private HF Space sandbox when explicitly selected; HF Jobs for remote compute
Evaluation Trackio metrics/alerts, job logs, saved outputs, one-pilot-before-batch rule, autonomous sweeps and ablations
Control Interactive approval events, scheduled-job hard stop, web-session cost caps, cancellation, session persistence, doom-loop checks
Notifications One-way Slack status/approval/error/completion notifications

Authority and privacy audit

Local execution is the default

The CLI's default tool_runtime is local, where bash, read, write, and edit act on the current machine. --sandbox-tools switches to HF Space tools; it does not make the default invocation safe. Run experiments in a scoped checkout or sandbox and give tokens only the permissions the task requires. Source: README, configs/cli_agent_config.json, and agent/tools/local_tools.py, reviewed 2026-08-10

A prompt means broad auto-approval

ml-intern "prompt" enters headless mode, sets config.yolo_mode = True, and auto-approves approval-gated actions except scheduled HF jobs. The CLI/headless legacy YOLO path is not constrained by the web session's billable cost cap. Normal use should be interactive with manual approvals; scheduled, billable, publishing, destructive, or durable mutations remain owner decisions. Source: agent/main.py and agent/core/agent_loop.py, reviewed 2026-08-10

Trace opt-out is narrower than it sounds

Defaults enable both save_sessions and share_traces. The personal copy is created private and best-effort scrubbed, and its dataset card explicitly warns that prompts, code, paths, repository names, task context, and tool outputs may remain. More importantly, share_traces: false disables only that personal copy. save_and_upload_detached() still sends a row containing scrubbed messages, events, tools, usage, and user_id to smolagents/ml-intern-sessions. That payload is broader than the README phrase “only receives anonymized telemetry rows.” There is no built-in local-save-only setting: save_sessions: false disables both persistence and upload. Source: agent/config.py, agent/core/session.py, and agent/core/session_uploader.py, reviewed 2026-08-10

This is the current adoption blocker for sensitive work. A reviewed wrapper or upstream option must separate local persistence from all network uploads before ML Intern becomes a global/default tool.

Research claims versus proof

The saved post reports a 10% to 32% GPQA improvement on Qwen3-1.7B in under ten hours, a 22.99% Claude Code comparison, a 60% HealthBench advantage over Codex, and a GRPO recovery through ablations. The original 16.58-second launch video visually progresses from an advertised 26.0% intermediate GPQA score to 39.0%. Those mismatched launch numbers are not necessarily false, but they are not a reproducible benchmark receipt. Source: X/@akseljoonas and preserved launch video, reviewed 2026-08-10

Do not upgrade them to benchmark facts without exact dataset/evaluator revisions, split and contamination checks, source/config/seeds, complete run logs, cost, selection policy, and an independent rerun. The tool's own workflow is valuable even while its marketing result remains unverified.

Kevin route

  • Recommend for a concrete HF-native research task that needs literature, dataset inspection, implementation, remote compute, and iterative evals.
  • Preserve, do not globally install until a concrete task exists and the no-upload/local-only boundary is fixed or wrapped.
  • Run interactively in sandbox mode, not headless/local YOLO, with scoped credentials and explicit approvals for compute, publishing, and mutations.
  • Use Research Experiment Workflow as the owner. ML Intern is one executor; it does not own the baseline, evaluator, provenance, or claim.
  • Require task-specific proof. The upstream 508 passing unit tests show implementation maturity, not ML result validity or safe live credentials.

Timeline

  • 2026-08-10 | Re-reviewed the complete source and launch video, preserved the upstream checkout/artifacts, passed 508 unit tests, extracted the durable research-experiment loop, and found the local-default, headless-YOLO, and shared-trace payload boundaries. Kept/recommended the capability but held global installation and sensitive use until local-only trace handling is proven. Source: X/@akseljoonas; huggingface/ml-intern; local capture manifest
  • 2026-07-02 | First deep review recorded the untagged 0.1.0 repository, Hugging Face and local-model routes, and launch-media benchmark caveat. Source: GitHub huggingface/ml-intern; local thumbnail
  • 2026-04-21 | @akseljoonas announced ML Intern as an open-source version of the Hugging Face post-training research loop. The saved source snapshot has 4,657 likes and 6,014 bookmarks. Source: X/@akseljoonas