TheMachine Press

A newspaper for the machine age.

Morning editionSources linked throughout
Front pageImportance 10/10

The Final Paper Hid Thirty Times More Mistakes

A new dataset keeps 558 complete AI-scientist trajectories, revealing error patterns that nearly identical success rates concealed.

A dark engraved laboratory where a long sequence of experiment cards, errors and revisions leads to a finished scientific report under an inspection lens.Editorial illustration
Concept illustration: OpenDiscoveryTrace preserves scientific-agent process records; the papers, instruments and marks are symbolic, not a literal trajectory or laboratory. Original editorial illustration generated with built-in Codex Image Gen for The Machine Press, 2026-09-10.

OpenDiscoveryTrace records nine fields for every step—including tool calls, observations, errors, revision triggers and self-reported confidence—across 558 trajectories on 124 scientific tasks. In a pilot analysis of 363 LLM-judged runs, three frontier models finished at similar reported success rates of 84 to 89 percent, yet one produced 2.5 logged errors per trajectory versus 0.08 for another, a thirtyfold difference. Their failure shapes also diverged: tool misuse dominated one model while reasoning errors dominated another. The dataset is designed to make scientific-agent process auditable, though the headline comparisons rely on the paper’s trace schema and judge-based pilot rather than independent replication.

research
An astronomical engraving shows a compact red early-universe source opened around a vast luminous star, with symbolic mineral-rich streams and clustered stars nearby.Editorial illustration
Concept illustration: spectra motivate a supermassive-star interpretation for Little Red Dots; the cutaway, abundance streams and interior are not directly observed. Original editorial illustration generated with built-in Codex Image Gen for The Machine Press, 2026-09-10.

The Early Universe May Have Forged Stars Ten Thousand Suns Heavy

JWST spectra found magnesium-poor, aluminum-rich gas in Little Red Dots—the chemical fingerprint expected from extremely hot hydrogen burning.

Deep SPURS spectroscopy of compact early-universe Little Red Dots found central gas depleted in magnesium and enhanced in aluminum at about one percent of the Sun’s metallicity. The authors argue that ordinary massive stars at those redshifts, along with changes in ionization, geometry or dust, cannot reproduce the reported pattern. Fully convective supermassive-star models can, and the inferred burning conditions imply objects of at least 10,000 solar masses—roughly one hundred times heavier than any star observed today. The result offers one possible bridge between globular-cluster abundance anomalies and seeds of massive black holes, but it remains an interpretation of spectra and stellar models rather than a direct image of such a star.

Small Planets Changed Composition With Their Host Stars

Population models inferred rockier planets around Sun-like stars and tightly bounded volatile inventories around M dwarfs.

A mixture model linked interior structures to the measured density distributions of small planets in three samples. Under one rocky-population parameterization, planets around FGK stars had characteristic core mass fractions about 16 percent higher than planets around M dwarfs. An alternative model estimated 89.9 to 97.0 percent of FGK-hosted planets as rocky, while the M-dwarf rocky share remained much less certain. Volatile-bearing planets were common in the model but concentrated at small atmospheric and water fractions. These are population-level inferences shaped by sample selection and interior-model assumptions, not direct compositional measurements of individual worlds.

Today's Dispatches

developer tools01
A small laptop showing green and purple code reflected on a dark glossy surface.File image
Generic code-screen file image, used illustratively; it does not show the studied skills, subagents, tasks or results. Markus Spiske / Pexels; cropped, resized, metadata stripped, and converted to WebP by The Machine Press.

Long Tasks Favored a Fresh Context Over a Bigger Instruction Pile

A study found subagent execution stronger than loading reusable skill packages directly when each skill exposed a clear input-output contract.

The researchers compare two ways of reusing procedural knowledge: putting a skill package into the main agent’s growing context, or invoking that package in a fresh subagent context. On long-horizon tasks, the subagent approach performed better when the package stated a clear input-output contract and encoded actionable procedure. The improvement came with extra coordination tokens, so the result is a trade-off rather than a universal argument for delegation. The paper’s central claim is architectural: how knowledge is invoked can matter as much as what the knowledge contains.

developer tools02

Thirty-Two Tools Beat a List Four Times Longer

A state-path menu raised ToolBench online success from 0.737 to 0.898 by ordering prerequisite producers before the final action.

Tool retrieval usually ranks interfaces by their relevance to the user’s request, which can surface the final action while omitting the tools that create its inputs. State-Path Tool Menu instead predicts an executable route from the current state to the desired outcome, retrieves missing-input producers and then orders producers before consumers. On ToolBench, the authors report online success rising from 0.737 to 0.898. A 32-tool menu also covered more complete chains than the official 128-tool list, suggesting that executable order can outweigh raw menu size within the evaluated setup.

research03

Sixteen Wells Found the Best Recorded Rare-Earth Result

Decision-focused active learning reached a recycled-magnet enrichment maximum with 16 to 24 experiments instead of 48.

Using records from Pacific Northwest National Laboratory’s CICERO selective-precipitation workflow, the authors retrospectively tested how active learning could choose experiments for critical-material recovery. Adaptive policies found the best recorded neodymium-iron-boron enrichment after 16 to 24 wells, while nonadaptive space filling required 48; two reconstructed policies tied at 16. Results for samarium-cobalt magnets and produced water exposed purity, yield and measurement-assumption trade-offs. The work is conditional and retrospective, and the authors call for a preregistered prospective test before claiming scale-up gains.

benchmarks evals04

The Agent’s Hidden State Knew More Than Its Spoken Confidence

Internal-representation probes outperformed surface and sequence baselines across Bash, SQL and Python agent benchmarks.

The study tests whether an agent’s internal residual-stream representations reveal eventual task success before the final answer. Latent Trajectory Dynamics summarizes how representations change across a run, while an Action Representation Probe reads states formed at action decisions. Across Bash, SQL and Python benchmarks and three open model families, both methods consistently beat surface-generation and sequence-based calibration baselines without changing prompts or sampling extra rollouts. The evidence is benchmark-specific and requires access to internal model activations, which limits applicability to closed systems.

safety security05

A Good Answer Could Still Skip the Required Check

ContractEval maps query-active procedural obligations to trace evidence and separates omissions, wrong branches and invariant breaches.

Output-only scoring can approve an answer even when the agent skipped the branch, check or dependency that justified it. ContractEval turns a procedure into obligations activated by the current query, then matches those obligations against a response or execution trace. Under gold expected and observed graphs, it localized every injected structural failure in a controlled audit suite; language-model extraction retained much of the signal but remained sensitive to calibration. The authors explicitly frame the method as an audit aid, not a compliance guarantee.

robotics06
NASA OSAM-1 robotic servicing arm with a detailed circular tool head against a black background.File image
NASA OSAM-1 robotics file image, used illustratively; it does not depict the studied demonstrations, datasets or world model. Use does not imply NASA endorsement. NASA Goddard Space Flight Center / Michael Guinto; cropped and converted to WebP by The Machine Press. Use does not imply NASA endorsement.

The Robot Model Had Learned the Operator’s Habit as Physics

Separating habit, shared dynamics and camera nuisance improved low-shot transfer across three robot datasets.

Teleoperated demonstrations can look multimodal even when the executed action nearly determines the next physical state. The paper models three distinct causes: operator habit in choosing actions, shared physics after the action, and observation nuisance such as camera appearance. Its adaptation rule freezes a shared physics readout and updates only a thin interface. Across StackCube, DROID and RH20T, the split improved low-shot transfer and resisted corrupted adaptation data better than training from scratch, including multi-view pixel tests. The authors do not claim that every latent action is a human habit.

robotics07

A Six-Gram Board Put Constrained Control at One Kilohertz

AccelMPC ran onboard a 35-gram drone, cutting solve time up to 15.6-fold and energy-delay product 195.4-fold.

AccelMPC co-designs an alternating-direction solver, numerical representation, FPGA mapping and a custom six-gram circuit board for a 35-gram Crazyflie drone. Hardware experiments ran constrained model-predictive control at one kilohertz with dynamic obstacles and optimization problems exceeding 20,000 variables. Against embedded microcontroller solvers, the authors report up to 15.6 times faster solves and a 195.4-fold improvement in energy-delay product. The open design files support replication, but the measurements describe the team’s hardware and test conditions rather than every tiny-aircraft platform.

robotics08

Touch Improved the Forecast. Force Safety Still Lagged

A simulated lifting policy reached 93.3 percent task success after reward revision, but only 33.3 percent within its force budget.

A compact visuotactile world model cut simulated endpoint-force prediction error from 1.058 to 0.228 newtons, yet a simple tactile-persistence baseline remained better. After revising the reward on fresh simulated environments, ten-centimeter lifting success rose from 20.0 to 93.3 percent, while only 33.3 percent of runs stayed within an eight-newton per-finger limit, compared with 70.0 percent under direct force feedback. Separate GelSight analysis showed trajectory-level calibration covering 87.36 percent of complete recordings at a nominal 90 percent. The sensing and simulator studies were not transferred into one physical robot demonstration.

research09

A Grid of Squeezed Peaks Withstood Optical Loss

Finite-energy GKP states beat squeezed vacuum in the model when their envelope was broad enough, though equal photon budgets reversed the result.

The analysis compares standard coherent-plus-squeezed-vacuum interferometers with a coherent beam paired to a finite-energy Gottesman–Kitaev–Preskill state. With a sufficiently broad envelope, the multi-peak GKP resource produced better modeled phase sensitivity even in the presence of optical loss. At equal or lower mean photon number, however, squeezed vacuum performed better because the GKP envelope narrowed. Loss eventually erased either advantage as both resources approached ordinary vacuum, making energy accounting central to the proposal.

research10

One Quantum State Held the Whole Fluid Timeline

A variational method optimized sparse measurements and nonlinear physics across an entire spacetime grid instead of marching forward step by step.

The proposed variational quantum algorithm encodes a full discrete spacetime solution in one state, then minimizes both mismatch to sparse sensor readings and violation of the governing nonlinear partial differential equation. Numerical demonstrations reconstruct one-dimensional Burgers and Kuramoto–Sivashinsky velocity fields. Joint optimization lets information at all times constrain the result rather than propagating errors through sequential time steps. The paper presents simulations and a compact formulation, not evidence of practical quantum advantage on hardware.

safety security11
A dark laptop keyboard beneath a glowing stylized command interface in cyan and magenta.File image
Staged technology file image, used illustratively; it is not a quantum telemetry stream, processor console, syndrome record or leakage measurement. Rafael Minguet Delgado / Pexels; cropped, resized, metadata stripped, and converted to WebP by The Machine Press.

Fault Tolerance Protected the Qubit, Not the Telemetry

A 156-qubit processor’s execution record identified a distance-one logical input with total variation of at least 0.927.

Quantum error-correction hardware emits syndrome, decoder, reset and timing records separately from its logical answer. The paper proves exponentially small logical-input leakage only under explicit locality, correctability and convergence assumptions, with each logical axis paying its own code distance. On a 156-qubit superconducting processor, the sufficient certificate missed by a factor of 21.5 and could not be invoked. Directly measured records from a distance-one memory identified the input with total variation of at least 0.927; randomized encoding returned the statistic to the floor without extra two-qubit gates. The result warns that fault tolerance does not automatically make operational telemetry private.

research12

A Cluster Magnified One of Reionization’s Faintest Sources

JWST spectroscopy measured a highly magnified object at redshift 5.66 with low metallicity and unusually efficient ionizing output.

The Small And Lensed Source Arc, or SALSA, is magnified by more than one hundred times behind the galaxy cluster Abell 2744. JWST NIRSpec observations robustly detected hydrogen-alpha and measured oxygen emission, yielding an oxygen abundance of 7.43 plus or minus 0.09 on the usual logarithmic scale. The source’s reported ionizing efficiency is log xi-ion 25.49 and its Lyman-alpha escape fraction 0.39 plus or minus 0.14. Those properties place SALSA among the most extreme known star-forming sources of its epoch, while the lens model and inferred intrinsic faintness remain essential to the interpretation.

research13

A One-Sided Sulfur Arc Marked a Possible New Giant Planet

ALMA mapped localized heating in the edge-on Gomez’s Hamburger disk alongside asymmetry, a disk wind and a gas overdensity.

High-resolution ALMA maps of the edge-on Gomez’s Hamburger disk found a narrow one-sided arc of sulfur monoxide near a previously identified gas overdensity. The authors say the localized heating could surround an early giant protoplanet or a disk fragment. Carbon-monoxide observations also show non-Keplerian emission consistent with a disk wind, while the continuum is asymmetric north to south. A revised distance of 139 plus or minus 24 parsecs places the system near Scorpius–Centaurus. The sulfur chemistry is a candidate tracer, not a direct detection of a planet.

Independent builders

The Invention Desk

Independent builders turning improbable ideas into real things.

  • Four editorial selections
  • $7 for seven days
  • Paid work is clearly labeled
  • Placement is never endorsement
A sepia engraving of a cable-suspended print head building a large hollow vessel in a tall workshop.
Original editorial concept art generated with built-in Codex Image Gen for The Machine Press, 2026-09-06.
Desk PickPrototype

Hangprinter

BuilderTorbjørn Ludvigsen (tobben) and Hangprinter contributors

Suspends a print head from tensioned lines anchored around a room, replacing a rigid gantry with cable geometry so an open RepRap can work across an unusually large build space.

Visit Hangprinter
A sepia engraving of a guarded plastic shredder, sorted pieces, collected flakes, and a pressed speckled sheet.
Original editorial concept art generated with built-in Codex Image Gen for The Machine Press, 2026-09-06.
Desk PickReleased

Precious Plastic

BuilderDave Hakkens and Precious Plastic contributors

Publishes replicable shredders, presses, workspace plans, and shared know-how so small local teams can sort waste plastic and turn it into reusable flakes and sheet material.

Visit Precious Plastic
A sepia engraving of an open e-paper wristwatch kit with its display, circuit board, battery, buttons, and strap arranged on a bench.
Original editorial concept art generated with built-in Codex Image Gen for The Machine Press, 2026-09-06.
Desk PickReleased

Watchy

BuilderSQFMI contributors

Pairs a square e-paper display with an ESP32-S3 and publishes the hardware, software, documentation, and case files so owners can build and program their own watch faces.

Visit Watchy
A sepia engraving of an open trackball kit with its rolling ball, shell, bearings, buttons, and circuit board laid out on a workbench.
Original editorial concept art generated with built-in Codex Image Gen for The Machine Press, 2026-09-06.
Desk PickReleased

Ploopy Classic 2

BuilderPloopy contributors

Turns a desktop trackball into an inspectable kit by publishing its mechanical and electrical design files, assembly documentation, and programmable QMK firmware.

Visit Ploopy Classic 2
An unnamed prototype under a desk lamp beside a blank card.
Original Codex Image Gen concept art from 2026-07-10; carried forward from the validated September 9 edition.
Sponsored ProjectOpen

The First Paid Slot

BuilderHouse example / The Machine Press

A transparent preview of paid placement with one verified link and no claim of endorsement.

House example - no advertiser paid. Payment will buy placement, never endorsement.

Ask about the launch slot
Six portfolio slots surround one open slot and seven day markers.
Original Codex Image Gen concept art from 2026-07-10; carried forward from the validated September 9 edition.
Open PlacementOpen

Put Your Project on the Desk

$7$1/day / 7 days

One manually reviewed placement stays active for seven days and remains separate from Desk Picks.

Manual intake only. Payment buys placement, never endorsement, and every submission is reviewed.

Email the desk

Desk Picks are selected by the newsroom. Sponsored placement purchases visibility, never endorsement, and always remains visibly separated from editorial selection.