# Agent Native Research (ARA Lab) > Agent Native Research — also known as ARA Lab, ARA Labs, and published at > https://www.agenticresearch.sh/ — builds the infrastructure layer for AI scientists. It > authored "The Last Human-Written Paper: Agent-Native Research Artifacts" > (arXiv:2604.24658) and stewards the open Agent-Native Research Artifact > (ARA) standard: an executable, verifiable format for research that both > humans and AI agents can audit, reproduce, and extend. ## What this lab is Agent Native Research is a research lab and open standards body working on science for AI rather than AI for science. AI for science applies a model to a scientific problem and improves one researcher at a time; science for AI rebuilds the practice of research so autonomous AI scientists can participate in it — the artifact format, the verification layer, the review process, and the incentives. Making each node ten times smarter while leaving the network untouched does not give ten times the science. The lab is unrelated to the Advanced Robotics and Automation (ARA) Laboratory, which shares the acronym. ## Agent-Native Research Artifact (ARA) An ARA is the unit of research in the AI-native ecosystem: an executable, verifiable knowledge package with four layers — scientific logic, runnable code with full specifications, an exploration graph that preserves the dead ends, and evidence grounding every claim in raw output. On PaperBench and RE-Bench, ARA raises question-answering accuracy from 72.4% to 93.7% and reproduction success from 57.4% to 64.4%. Publishing one takes about 15 minutes: run the `/submit-ara` agent skill, which compiles a research directory into a valid ARA, pushes it to your public GitHub account, and registers it in the ARA Hub as an interactive step-by-step replay. ## Primary sources - [The Last Human-Written Paper: Agent-Native Research Artifacts](https://arxiv.org/abs/2604.24658): the paper that introduced the ARA standard (arXiv:2604.24658) - [The Last Human-Written Paper (essay)](https://www.agenticresearch.sh/blog/the-last-human-written-paper): the argument in full, published by the lab - [ARA open standard](https://github.com/ARA-Labs/Agent-Native-Research-Artifact): the specification, reference tooling, and agent skills - [ARA Hub](https://www.agenticresearch.sh/hub): published artifacts, replayable step by step - [Research](https://www.agenticresearch.sh/research): EXP-Bench, Sci-Reasoning, Curie, and the ARA protocol work - [Submit an ARA](https://www.agenticresearch.sh/submit): how to publish a research artifact to the Hub ## Field reports - [AI Should Build Its Own Research World Model](https://www.agenticresearch.sh/blog/research-world-model): Opus 4.8 Cleared ARC-AGI’s Final Level in One Shot. - [The Second Half of AI for Science](https://www.agenticresearch.sh/blog/the-second-half-of-ai-for-science): Make the node 10× smarter and leave the network untouched — you don't get 10× science. The second half is about rebuilding the ecosystem, not the scientist. - [The Last Human-Written Paper](https://www.agenticresearch.sh/blog/the-last-human-written-paper): When neither the author nor the audience is human, the three-century-old paper format stops making sense. ## Published artifacts - [NanoGPT Speedrun](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/speedrun/nanogpt-speedrun) - [Restricted-Architecture MLM (RE-Bench task)](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/rebench/rebench-restricted_mlm) - [Rust CodeContests Inference (RE-Bench task)](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/rebench/rebench-rust_codecontests) - [Triton Cumsum Kernel (RE-Bench task)](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/rebench/rebench-triton_cumsum) - [APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/adaptive-pruning) - [All-in-one simulation-based inference](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/all-in-one) - [Batch and Match: Black-Box Variational Inference with a Score-Based Divergence](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/bam) - [BBOX-ADAPTER: Lightweight Adapting for Black-Box Large Language Models](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/bbox) - [Andes: Defining and Enhancing Quality-of-Experience in LLM-Based Text Streaming Services](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/extra/andes) - [EXP-Bench: Can AI Conduct AI Research Experiments?](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/extra/expbench) - [Venn: Resource Management for Collaborative Learning Jobs](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/extra/venn) - [Efficient Transfer Learning in Diffusion Models via Adversarial Noise](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/bridging-data-gaps) - [Unsupervised Zero-Shot Reinforcement Learning via Functional Reward Encodings](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/fre) - [Fine-tuning Reinforcement Learning Models is Secretly a Forgetting Mitigation Problem](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/ftrl) - [Refined Coreset Selection: Towards Minimal Coreset Size under Model Performance Constraints](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/lbcs) - [LCA-on-the-Line: Benchmarking Out-of-Distribution Generalization with Class Taxonomies](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/lca-on-the-line) - [A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/mechanistic-understanding) - [Challenges in Training PINNs: A Loss Landscape Perspective](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/pinn) - [RICE: Breaking Through the Training Bottlenecks of Reinforcement Learning with Explanation](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/rice) - [Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/robust-clip) - [Sample-specific Masks for Visual Reprogramming-based Prompting](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/sample-specific-masks) - [SAPG: Split and Aggregate Policy Gradients](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/sapg) - [Self-Composing Policies for Scalable Continual Reinforcement Learning](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/self-composing-policies) - [Self-Expansion of Pre-trained Models with Mixture of Adapters for Continual Learning](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/self-expansion) - [Semantic Self-Consistency: Enhancing Language Model Reasoning via Semantic Weighting](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/semantic-self-consistency) - [Sequential Neural Score Estimation](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/sequential-neural-score-estimation) - [Stay on topic with Classifier-Free Guidance](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/stay-on-topic-with-classifier-free-guidance) - [Stochastic Interpolants with Data-Dependent Couplings](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/stochastic-interpolants) - [Test-Time Model Adaptation with Only Forward Passes](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/test-time-model-adaptation) - [What Will My Model Forget? Forecasting Forgotten Examples in Language Model Refinement](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/paperbench/what-will-my-model-forget) - [Fix Embedding (RE-Bench task)](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/rebench/rebench-fix_embedding) - [nanoGPT Chat RL (RE-Bench task)](https://www.agenticresearch.sh/ara/AmberLJC/ara-paperbench/artifacts/rebench/rebench-nanogpt_chat_rl) - [World-Model ARA for ARC-AGI-3 ls20 (Locksmith)](https://www.agenticresearch.sh/ara/ARA-Labs/ara-ls20): An agent infers all mechanics and win conditions of the ARC-AGI-3 Locksmith game purely from action-diff observation, solving all seven levels including a fog-gated final level. - [Understanding Agent Performance on PostTrainBench](https://www.agenticresearch.sh/ara/AmberLJC/ara-posttrainbench-agent-performance): Across 1,226 post-training agent runs, performance is ~90% determined by agent and task identity rather than execution, and reward hacking is real but does not improve scores. - [ARA Demo](https://www.agenticresearch.sh/ara/ARA-Labs/ARA-Demo): A Codex autonomous agent reduced the 124M-GPT step count from 3500 to 2949 across four optimizer-search waves, with a novelty wave yielding a clean negative result and a compliance quarantine reshaping the v2 frontier. ## Also known as ARA, ARA Lab, ARA Labs, ARA Commons, Agentic Research, Evolving Lab, EvolvingLab, Agent Native Research Lab, AI-Native Research. ## Contact amber@ara-commons.com · https://x.com/ainativescience · https://github.com/ARA-Labs