Paper Feed

Issue 23 · Jun 1–7, 2026

Every candidate

All 3,877 papers were scored from their abstracts by gpt-5.6-luna; 1,444 were not skipped. Shown here: the top 150 of those, in score order. Picks are marked.

  1. strong AI / ML picked▲ 4 score 5.9

    Frames2LoRA: Parametric Video Internalization for Vision-Language Models

    Manan Suri, Sarvesh Baskar, Dinesh Manocha

    Frames2LoRA uses a hypernetwork to convert a video’s intermediate VLM representations directly into a LoRA adapter, so the model can answer later queries without retaining video tokens in context. On SmolVLM2 models, it reportedly matches direct video-context inference across most captioning and QA evaluations, while cutting query-time visual-token load by up to 1,500× and latency by 6–80×; independently generated adapters can also compose across video segments.

    This is a genuinely new way to internalize video into model parameters and reports an unusually large efficiency gain while preserving performance, including stability on much longer and higher-resolution videos.

  2. strong AI / ML picked score 5.6

    Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities

    Yousef Radwan

    The paper proposes finding a sentiment/valence direction in a frozen encoder using only nine emotion categories and short example narratives, rather than many supervised sentiment labels. It reports that related directions transfer across text, images, audio, and EEG, support zero-shot valence prediction, and causally affect LLM sentiment behavior when ablated; the effect is strong for continuous valence but not categorical concepts and varies across model families.

    The combination of extremely weak supervision, cross-modality transfer into independently trained encoders and brain recordings, and model-specific causal ablations would be a genuinely important result if the alignment and controls hold up.

  3. strong Robotics picked score 5.4

    Wave Focusing in Metamaterials: Tactile Displays Beyond the Diffraction Limit

    Gregory Reardon, Max Linnander, Dustin Goetz et al.

    The paper uses a lattice of mechanical resonators embedded in a flexural plate to alter wave propagation and focus vibrations more tightly than an ordinary plate allows. A fabricated display achieves a tenfold reduction in virtual-pixel area and supports independently controlled, perceptually localized single- and multi-point sensations, including moving tactile sources, with relatively few actuators.

    It demonstrates a genuinely new hardware/physics approach to overcoming the diffraction limit in distributed haptic displays, backed by fabrication and behavioral experiments rather than simulation alone.

  4. strong Robotics score 5.4

    On the Hardness of Optimal Motion on Trees

    Tzvika Geft

    The paper proves that optimizing multi-agent motion remains NP-hard even on trees, covering labeled and two-colored agents under distance, makespan, and flowtime objectives. It also resolves the longstanding complexity of optimal Pebble Motion on trees and proves the first hardness result for colored Pebble Motion on any graph class, via a new NP-hardness result for Stack Rearrangement; remarkably, subdivided stars already suffice.

    It settles decades-old open complexity questions and exposes a common hardness barrier across several basic motion-planning models, with hardness holding on extremely simple tree structures.

  5. maybe AI / ML ▲ 60 score 5.2

    Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses

    Pengcheng Jiang, Zhiyi Shi, Kelly Hong et al.

    Harness-1 trains a 20B retrieval agent with reinforcement learning while moving routine bookkeeping out of the model and into an explicit search harness. The harness tracks candidates, evidence, verification, deduplication, and context budgets, leaving the policy to make semantic search and stopping decisions. Across eight web, finance, patent, and multi-hop QA benchmarks, it reports 0.730 average curated recall—11.4 points above the strongest open search subagent—and particularly strong transfer to held-out domains.

    Externalizing search state during RL is a meaningful and potentially generalizable agent-training direction, with a substantial reported gain and cross-domain transfer, but the abstract does not establish how much is due to the specific harness components or how competitive the comparisons really are.

  6. strong Robotics picked▲ 7 score 5.2

    Flash-WAM: Modality-Aware Distillation for World Action Models

    Arman Akbari, Ci Zhang, Arash Akbari et al.

    Flash-WAM distills a joint video-and-action diffusion model into one denoising step per modality, using different consistency parameterizations to account for their different noise regimes. On LingBot-VA, it cuts RoboTwin inference latency from 8.1 seconds to 348 ms while retaining 85.5% RoboTwin and 95.7% LIBERO success; on a Unitree G1, it reaches 60% average success versus 24% for naive distillation.

    The modality-specific distillation design addresses a real obstacle in joint video-action models, and the reported 23x speedup plus substantial real-robot recovery makes it more than a routine compression result.

  7. strong AI / ML score 5.2

    Sharp First-Order Lower Bounds for Higher-Order Smooth Nonconvex Optimization

    Dongruo Zhou

    This paper proves matching first-order oracle lower bounds for finding stationary points in nonconvex optimization when the objective has higher-order smoothness. In particular, it shows that the known upper-bound rates of ε^-7/4 for Lipschitz Hessians and ε^-5/3 for Lipschitz third derivatives are optimal, using a dimension-free block-chain hard instance that extends to any finite smoothness order.

    It closes a longstanding lower-bound gap and establishes optimality of accelerated first-order rates across higher-order smoothness classes, rather than offering another algorithmic variant.

  8. maybe AI / ML ▲ 70 score 5.2

    Evolving Agents in the Dark: Retrospective Harness Optimization via Self-Preference

    Wenbo Pan, Shujie Liu, Chin-Yew Lin et al.

    The paper proposes Retrospective Harness Optimization, which improves an agent’s tools, skills, and workflows using only its own past trajectories rather than labeled validation data. It selects difficult prior tasks, re-solves them, generates candidate harness changes, and chooses among them using self-validation and pairwise self-preference. The authors report a large SWE-Bench Pro pass-rate increase from 59% to 78% after one round, along with gains in technical and knowledge-work settings and better behavior over long sessions.

    The combination of retrospective failure targeting and self-preference for harness updates is a potentially important route to unlabeled agent improvement, and the reported 19-point SWE-Bench gain is substantial, but the abstract does not establish how reliable self-evaluation is or provide enough detail to validate the broad claims.

  9. maybe AI / ML picked▲ 30 score 5.1

    Streaming Communication in Multi-Agent Reasoning

    Zhen Yang, Xiaogang Xu, Wen Wang et al.

    StreamMA lets downstream agents consume each intermediate reasoning step as soon as it is produced, rather than waiting for an upstream agent to finish its whole chain. The authors report lower latency and, unexpectedly, better accuracy, arguing that early reasoning steps are more reliable and that exposing later steps can propagate errors; across eight benchmarks, two frontier models, and three multi-agent topologies, it improves average performance by 7.3 percentage points.

    The combination of streaming communication and the claim that partial, earlier reasoning can outperform complete chains is a genuinely interesting systems and reasoning insight, with unusually broad reported evaluations, but the abstract does not establish how robust the gains are or whether the latency/cost tradeoffs hold in practice.

  10. maybe AI / ML ▲ 28 score 5.0

    Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation

    Yuxuan Bian, Zeyue Xue, Songchun Zhang et al.

    Echo-Infinity is an autoregressive video generator that replaces fixed KV-cache eviction and compression rules with learnable memory queries. These queries are updated as old frames leave the local context, giving constant-cost history compression, while a revised relative-RoPE scheme avoids the model’s finite temporal-position limit. The authors report state-of-the-art short- and long-video results and a real-time rollout exceeding 1.3 million frames (about 24 hours), though the abstract gives few quantitative details.

    The learnable, evolving memory plus a practical solution to unbounded temporal RoPE is a meaningful direction, and the claimed 24-hour real-time rollout would be important, but the abstract provides insufficient quantitative or comparative evidence to justify a strong verdict.

  11. maybe AI / ML ▲ 126 score 5.0

    Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models

    Mahtab Bigverdi, Linjie Li, Weikai Huang et al.

    The paper trains a vision-language model to produce intermediate “imaginative perception” tokens representing what would be observed from unseen viewpoints or through occlusions, rather than forcing this reasoning entirely into text. On perspective-taking, path-tracing, and multiview-counting tasks, this supervision improves performance; the largest reported gain is 3.4% on multiview counting, and it can outperform textual chain-of-thought training in these spatial settings.

    The modality-specific intermediate representation and finding that textual chain-of-thought can hurt spatial reasoning are genuinely interesting, but the reported gains and evidence base appear too limited for a stronger verdict.

  12. maybe AI / ML ▲ 22 score 4.9

    OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents

    Rui Yang, Qianhui Wu, Yuxi Chen et al.

    OpenWebRL trains visual web agents with online, multi-turn reinforcement learning directly on live websites rather than relying mainly on large static demonstration datasets. Its 4B model reportedly reaches 67.0% on Online-Mind2Web and 64.0% on DeepShop using only 0.4K initialization trajectories and 2.2K RL tasks, while the paper analyzes the infrastructure, judging, context handling, and optimization choices involved.

    The combination of live-browser online RL, relatively little supervised data, and competitive open-agent performance is a meaningful direction, but the abstract does not provide enough detail about evaluation breadth, cost, baseline comparability, or how much of the gains come from engineering choices to warrant a strong verdict.

  13. maybe AI / ML ▲ 22 score 4.9

    Bootstrap Your Generator: Unpaired Visual Editing with Flow Matching

    Yoad Tewel, Yuval Atzmon, Gal Chechik et al.

    Bootstrap Your Generator trains image- and video-editing flow-matching models without paired before/after examples or external reward models. It uses semantic editing cues from a frozen base model, cycle consistency, and a gradient-routing technique that transfers losses from clean predictions to noisy training states. The authors report better results than supervised baselines trained on millions of paired samples, including on unseen domains, though the abstract gives no quantitative details.

    The combination of unpaired training, self-derived semantic supervision, and gradient routing could substantially reduce the data cost of image/video editing, but the abstract provides insufficient quantitative evidence to warrant a strong recommendation.

  14. maybe AI / ML score 4.9

    Building The Ph(ysical)AI Layer Of Machine Intelligence

    Ulbert Jose Botero, Liam Smith, Brooks Olney et al.

    The paper proposes a 1.99M-parameter encoder trained only on radio-frequency data, with Fourier, energy-conservation, and symmetry constraints built into its architecture and losses. Without fine-tuning, frozen representations reportedly transfer via linear probes to 15 tasks across RF, audio, images, text, and video, reaching 77.7% average accuracy, with better results on physically grounded tasks than semantic ones. The claimed contribution is that physical principles, rather than massive multimodal data, can support some cross-modal generalization while exposing a boundary between physical and semantic understanding.

    The RF-only, principle-driven cross-modal transfer claim is genuinely unusual and potentially important, but the abstract does not provide enough baseline, dataset, or evaluation detail to establish that the impressive aggregate results are broadly meaningful rather than task-selection or representation-transfer artifacts.

  15. strong Robotics score 4.9

    Learning All-Terrain Locomotion for a Planetary Rover with Actively Articulated Suspension

    Arthur Bouton, Tristan D. Hasseler, Michael Paton et al.

    The paper presents ERNEST, a four-wheeled rover with actively actuated suspension that can reconfigure its wheels, steer, and redistribute loads. A unified reinforcement-learning controller, trained in a terramechanics simulator and consolidated across terrain-specialized policies, transfers to a physical rover without explicit terrain classification; it traverses rocks, steps, traps, ripples, and slopes, reducing cost of transport by 37% on dry sand and remaining mobile on wet sand where passive suspension fails.

    This is a credible real-robot demonstration of learned, unified control exploiting actively articulated suspension across substantially different terrains, including a striking wet-sand failure case for passive suspension, though the abstract does not establish how broadly the result generalizes.

  16. maybe Neuroscience picked score 4.9

    Predictable Mean-Field Chaos in Random Recurrent Neural Networks

    Alkesh Yadav, Vladimir Shaidurov, Jonathan Kadmon

    This paper argues that a random recurrent network can have positive Lyapunov exponents yet be perfectly predictable from the exact continuous-time history of a single neuron. Using dynamical mean-field theory and a Lanczos/Krylov representation, it links this predictability to the smoothness of the nonlinearity and shows that finite temporal modes give a prediction horizon growing logarithmically with the number of modes. The main conceptual result is that microscopic chaos and the amount of recoverable predictive information can scale differently.

    The claim that chaotic mean-field activity can be deterministically recoverable from one neuron's past, despite positive Lyapunov instability, is genuinely surprising and relevant to how neural variability is interpreted, but the abstract provides limited empirical or finite-network validation.

  17. maybe Robotics ▲ 23 score 4.8

    NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation

    NVIDIA, :, Aarti Basant et al.

    OmniDreams adapts NVIDIA’s Cosmos diffusion model into a real-time, autoregressive, action-conditioned video simulator for closed-loop autonomous-driving evaluation. Trained on 21,000 hours of driving data, it generates sensor observations conditioned on history, simulator state, and the vehicle’s actions, including novel weather and agent behaviors; a smaller policy built from it reportedly outperforms Alpamayo 1.5 on a NuRec benchmark.

    The combination of real-time generative simulation, closed-loop action conditioning, and using the world model as a policy backbone is potentially important, but the abstract gives no quantitative results for realism, latency, long-horizon consistency, or policy gains beyond a preliminary benchmark claim.

  18. maybe AI / ML ▲ 96 score 4.8

    Code2LoRA: Hypernetwork-Generated Adapters for Code Language Models under Software Evolution

    Liliana Hotsko, Yinxi Li, Yuntian Deng et al.

    Code2LoRA uses a hypernetwork to generate repository-specific LoRA adapters from code, avoiding extra repository tokens at inference time. A GRU-based variant updates the adapter state from successive code diffs, and on a 604-repository benchmark it matches per-repository LoRA on static tasks and improves over a shared LoRA by 5.2 percentage points on evolving repositories.

    The combination of generated repository adapters and incremental diff-based adaptation is a meaningful alternative to long-context retrieval or costly per-repository fine-tuning, supported by broad benchmark results, but the reported gains do not yet establish a major capability or efficiency breakthrough.

  19. strong AI / ML picked score 4.8

    Goedel-Architect: Streamlining Formal Theorem Proving with Blueprint Generation and Refinement

    Jui-Hui Chung, Ziyang Cai, Zihao Li et al.

    Goedel-Architect organizes Lean theorem proving around a global dependency graph of definitions and lemmas, then proves graph nodes in parallel and revises the graph when attempts fail. Using a 284B-parameter DeepSeek model, it reports 99.2% pass@1 on MiniF2F-test, 75.6% on PutnamBench, and strong results on recent olympiad problems; seeding the graph with natural-language proofs improves the harder benchmarks further. The main new idea is replacing purely recursive proof decomposition with globally planned, failure-refined proof blueprints, alongside a claimed large cost reduction.

    The global blueprint-and-refinement formulation appears meaningfully different from standard recursive theorem decomposition, and the reported performance across multiple difficult formal-math benchmarks—including recent olympiad problems—suggests a substantial capability and efficiency advance, though the reliance on a very large backbone warrants verification.

  20. maybe AI / ML picked▲ 142 score 4.8

    Cosmos 3: Omnimodal World Models for Physical AI

    NVIDIA, :, Aditi et al.

    Cosmos 3 presents a unified mixture-of-transformers model that can take in and generate language, images, video, audio, and action sequences. NVIDIA claims it can serve as a common backbone for multimodal understanding, generation, simulation, and robot policies, and reports leading results on several public evaluations, with code, checkpoints, data, and benchmarks released openly.

    The breadth of a single open model spanning perception, generation, and action is potentially important, but the abstract gives no quantitative results or evidence that the unified approach materially outperforms specialized systems.

  21. maybe AI / ML ▲ 6 score 4.8

    AdaCodec: A Predictive Visual Code for Video MLLMs

    Haowen Hou, Zhen Huang, Zheming Liang et al.

    AdaCodec replaces repeated per-frame RGB tokens with a predictive video representation: it sends a full reference frame when prediction is difficult and otherwise encodes motion and residual changes as compact tokens. Against a Qwen3-VL-8B per-frame baseline at the same visual-token budget, it reportedly improves all 11 benchmarks; with one-seventh the budget, it exceeds the baseline on long-video tasks and reduces time-to-first-token from 9.26s to 1.62s on general-video benchmarks.

    The predictive visual-code interface is a meaningful, potentially general solution to temporal token redundancy, and the large reported efficiency gains across many benchmarks merit inspection, though the abstract does not establish how much comes from codec design versus implementation or evaluation choices.

  22. maybe AI / ML ▲ 17 score 4.8

    Complexity-Balanced Diffusion Splitting

    Noam Issachar, Dani Lischinski, Raanan Fattal

    The paper proposes splitting a diffusion model’s time range among specialized subnetworks, assigning larger or smaller segments according to estimated local approximation difficulty rather than using heuristic splits. It estimates this difficulty from flow energy or trajectory acceleration with a lightweight auxiliary model, and reports roughly a 35% FID improvement for CFG-equipped SiT-XL over naive temporal partitioning across several architectures and datasets, without increasing per-step inference cost.

    This is a principled approach to reducing the inefficiency of using one uniform-capacity network across diffusion time, with substantial reported gains, but the abstract does not establish how much extra training or inference complexity the partitioning and auxiliary modeling require.

  23. maybe AI / ML ▲ 25 score 4.8

    ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning

    Ziyan Liu, Xueda Shen, Yuzhe Gu et al.

    ThoughtFold trains reasoning models to remove unnecessary trial-and-error steps from otherwise correct chain-of-thought trajectories. It uses introspection to generate shorter candidate sub-trajectories and a masked preference objective to favor direct links between essential reasoning steps; on DeepSeek-R1-Distill-Qwen-7B, it reportedly cuts token use by about 56% while preserving accuracy.

    A large claimed inference-efficiency gain from fine-grained trajectory folding is worth checking, but the abstract gives limited detail about evaluation breadth, accuracy tradeoffs, and whether the method generalizes beyond one distilled model.

  24. maybe Robotics ▲ 42 score 4.8

    Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking

    Zekun Qi, Xuchuan Chen, Dairu Liu et al.

    Humanoid-GPT uses a causal Transformer pretrained on a 2-billion-frame corpus combining motion-capture data and in-house recordings for whole-body humanoid control. The authors claim that scaling data and model size enables one controller to track highly dynamic motions and generalize zero-shot to unseen motions and control tasks, rather than relying on task-specific shallow trackers.

    The combination of large-scale motion pretraining and a general-purpose generative humanoid controller is a notable direction, but the abstract gives no quantitative results, hardware scope, or details showing that the claimed zero-shot generalization is more than broad but conventional scaling.

  25. maybe AI / ML ▲ 21 score 4.7

    ZipSplat: Fewer Gaussians, Better Splats

    Alexander Veicht, Sunghwan Hong, Dániel Baráth et al.

    ZipSplat replaces the usual one-Gaussian-per-input-pixel representation with a compact set of clustered scene tokens, each decoding into a group of Gaussians. This lets one pose-free feed-forward model trade reconstruction quality for representation size at inference time; it reportedly uses about 6× fewer Gaussians while improving PSNR over prior baselines and generalizing to unseen datasets.

    The decoupling of Gaussian count from image resolution, combined with inference-time quality–efficiency control and sizable reported gains, is a meaningful advance, but the abstract lacks absolute metrics and detailed evidence needed for a strong recommendation.

  26. maybe AI / ML ▲ 7 score 4.7

    Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction

    Tianxiang Jiang, Linquan Wu, Sheng Xia et al.

    Future-L1 lets a video language model interleave ordinary text reasoning with continuous latent visual representations of possible future frames, rather than expressing all intermediate reasoning in words. Training uses future-frame embedding alignment and a latent-aware reinforcement-learning objective; on two video event-prediction benchmarks it reports large gains, including 61.0 to 85.4 for Qwen3-VL-8B on FutureBench and 2.44 to 3.04 on TwiFF-Bench.

    The interleaved latent-visual reasoning mechanism is a genuinely interesting direction and the reported gains are unusually large, but the abstract provides limited evidence about benchmark breadth, ablations, and whether the improvements generalize beyond these tasks.

  27. maybe AI / ML score 4.7

    Compile Once, Differentiate Everywhere: A Differentiable Meta-Circular Interpreter

    Lucas Sheneman

    The paper compiles a self-hosting subset of Scheme into an autodiff-compatible graph, so a single compiled interpreter can execute new programs supplied as data while differentiating through their continuous constants. It proves almost-everywhere gradient correctness and validates numerical agreement on 171 recursive and higher-order programs, then uses the system to jointly search over program structure and parameters in battery degradation and El Niño inverse-problem tasks.

    Differentiating through a compiled meta-circular interpreter is a genuinely unusual capability that could broaden program-and-parameter search, but the demonstrated applications are limited and the abstract does not yet establish a large practical advantage over competing differentiable-programming or neuro-symbolic systems.

  28. maybe AI / ML ▲ 29 score 4.7

    VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization

    Junhao Cheng, Liang Hou, Tianxiong Zhong et al.

    This paper uses a vision-language model as a process evaluator rather than a planner: it extracts task-specific constraints and turns them into differentiable rewards for online LoRA optimization of a video-generation reasoner at test time. On symbolic and general video-reasoning benchmarks, the method reportedly improves performance by 16.7 points on average, substantially exceeding textual VLM guidance and Best-of-N sampling at similar test-time cost.

    The teacher-based, reward-guided test-time adaptation framing is a meaningful departure from using VLMs as solvers, but the abstract gives limited detail about absolute results, optimization stability, and whether the gains generalize beyond the cited benchmarks.

  29. maybe AI / ML ▲ 29 score 4.7

    OpenSkill: Open-World Self-Evolution for LLM Agents

    Zhiling Yan, Dingjie Song, Hanrong Zhang et al.

    OpenSkill studies agents that must improve after deployment without curated skills, successful demonstrations, verifiers, or target-task supervision. It builds knowledge and verification anchors from public documentation, code, and web resources, then turns them into transferable skills and practice tasks; experiments on three benchmarks and two agents report the best automated pass rates under this constraint, though the abstract gives no quantitative results.

    The notable contribution is a concrete framework for bootstrapping both skills and verification signals from open-world resources rather than assuming an existing learning loop, but the abstract lacks enough quantitative detail to justify a stronger verdict.

  30. maybe AI / ML picked▲ 16 score 4.7

    dots.tts Technical Report

    Shi Lian, Changtao Li, Bohan Li et al.

    dots.tts is a 2B-parameter multilingual text-to-speech model that autoregressively generates speech in a continuous latent space. It combines a specially trained AudioVAE, full-history conditioning for flow matching, and reward-free self-correction, reporting low WER, strong speaker similarity, and first-packet latency as low as 54–85 ms; code and checkpoints are released under Apache 2.0.

    The combination of continuous autoregressive TTS, long-context flow conditioning, self-correction, and efficient MeanFlow distillation is a substantial engineering direction with unusually strong reported quality and latency, but the abstract gives limited comparative detail and mostly supports open-source SOTA claims without enough evidence for a strong verdict.

  31. maybe AI / ML ▲ 75 score 4.7

    On the Geometry of On-Policy Distillation

    Zhennan Shen, Yanshu Li, Qingyu Yin et al.

    This paper studies where on-policy distillation updates move in model parameter space, comparing them with supervised fine-tuning and verifiable-reward RL. It reports that OPD quickly concentrates its cumulative changes into a narrow, low-dimensional subspace; restricting training to that early subspace preserves OPD performance but hurts SFT, suggesting OPD has a distinct and functionally sufficient update geometry.

    The subspace-locking result and the claim that OPD is geometrically distinct from both SFT and RLVR are a meaningful mechanistic finding, but the abstract provides limited quantitative detail and does not yet establish broad practical consequences.

  32. maybe AI / ML ▲ 122 score 4.7

    Audio Interaction Model

    Zhifei Xie, Zihang Liu, Ze An et al.

    The paper proposes an always-on audio interaction model that continuously listens, tracks context, decides when to speak, and responds asynchronously rather than operating in turn-based or offline mode. It combines streaming data and supervision, dual-loss training, FIFO inference, and a large multi-task corpus, claiming improved robustness, long-stream interaction, and proactive responses while retaining standard audio-task performance.

    The continuous perceive–decide–respond formulation and proactive-intervention capability are relevant and plausibly novel for audio models, but the abstract gives no quantitative results or convincing evidence that the bundled training and infrastructure produce a substantial advance.

  33. maybe AI / ML ▲ 15 score 4.7

    Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders

    Georgii Aparin, Vadim Popov, Tasnima Sadekova et al.

    The paper finds that Whisper’s internal audio representations contain linearly separable signals indicating when the model is hallucinating text for non-speech input, with the signal becoming stronger in deeper layers and concentrated in sparse features. Steering either the hidden activations or Sparse Autoencoder latents substantially reduces hallucinations, lowering rates from 72.63% to 14.11% for Whisper small and from 86.88% to 27.33% for large-v3, while causing only a small reported WER increase on speech.

    The combination of internal hallucination detection and SAE-based inference-time steering is a meaningful, potentially generalizable technique with large reported reductions, but the abstract leaves robustness across conditions and the exact speech-quality tradeoff insufficiently documented for a strong verdict.

  34. maybe Robotics ▲ 8 score 4.6

    AFUN: Towards an Affordance Foundation Model for Functionality Understanding

    Zhaoning Wang, Yi Zhong, Jiawei Fu et al.

    AFUN predicts both where a robot should interact with an object and the subsequent 3D motion, conditioned on a single RGB-D view and a language instruction. It also introduces a unified affordance-data pipeline combining robot, human, simulation, and scan data; across eight segmentation test sets and several motion tests, it reports large gains over prior methods and demonstrates embodiment-independent real-robot deployment without task-specific heuristics.

    The joint prediction of functional contact regions and executable 3D post-contact motion, backed by broad cross-source evaluation, is a meaningful step toward general-purpose manipulation, but the abstract does not establish how broadly the real-world capability transfers or whether the gains reflect a genuinely new model rather than strong data and training engineering.

  35. maybe AI / ML ▲ 7 score 4.6

    Multi-Agent Computer Use

    Jing Yu Koh, Ruslan Salakhutdinov, Daniel Fried

    The paper replaces a single serial computer-use agent with a manager that builds and continually revises a dependency DAG, dispatching multiple subagents in parallel while preserving information that later agents cannot directly observe. Across desktop and web benchmarks, this reportedly improves task success by 3.4–25.5% and reduces wall-clock completion time by about 1.5× on a long-horizon benchmark. The main contribution is treating decomposition, parallel execution, and persistent intermediate state as a unified scaling strategy for computer-use agents.

    This is a credible and fairly general systems direction for improving long-horizon computer use, supported by results across several benchmarks, but the abstract does not establish a fundamental capability breakthrough or clearly separate the gains from increased inference-time compute.

  36. maybe AI / ML ▲ 16 score 4.6

    Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them

    Woojung Han, Seil Kang, Youngjun Jun et al.

    The paper reports that image-to-video diffusion outputs from only two denoising steps can preserve more physically plausible motion than standard 50-step generation. It attributes this to degradation of motion-related spectral phase during denoising, and proposes PhaseLock, which extracts the early motion prior and applies it during later refinement without retraining; the method reportedly improves physical-consistency scores by 6.2 points with little extra compute.

    The counterintuitive finding that longer diffusion refinement can erase useful motion structure, plus a cheap training-free correction, is genuinely interesting, but the abstract lacks benchmark details and independent evidence that the phase explanation generalizes across models.

  37. maybe Robotics ▲ 9 score 4.6

    GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors

    Tianyi Xie, Haotian Zhang, Jinhyung Park et al.

    GRAIL presents a fully virtual pipeline for generating humanoid loco-manipulation demonstrations. It combines known 3D scenes and objects with video-model priors to reconstruct metric human-object trajectories, retargets them to a humanoid, and trains visual policies; using over 20,000 generated sequences, the policies achieve 84% pickup success and 90% stair-climbing success on a Unitree G1.

    The notable contribution is a scalable, entirely virtual route to robot-compatible whole-body manipulation and locomotion data with real-robot validation, though the abstract gives limited comparisons and covers only a few task categories.

  38. maybe Robotics ▲ 9 score 4.6

    World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis

    Yi Yang, Zhihong Liu, Siqi Kou et al.

    The paper introduces a World-Language-Action model that jointly predicts language-level intentions, future visual states, and robot actions using an autoregressive Transformer. Its world-modeling branch is trained from egocentric and cross-embodiment videos, while meta-queries let the model use predicted states to improve action generation without requiring world prediction at deployment; the 2B-parameter prototype runs at 40 ms per inference and reports strong results in simulated and real-world tasks.

    The unified autoregressive treatment of semantic intent, physical dynamics, and action—and the possibility of learning across robot embodiments without action labels—are substantive ideas, but the abstract provides limited comparative detail and the strongest claims are not yet enough for a strong verdict.

  39. maybe AI / ML ▲ 8 score 4.5

    OPRD: On-Policy Representation Distillation

    Shenzhi Yang, Guangcheng Zhu, Bowen Song et al.

    The paper distills a teacher language model into a student using hidden-state matching on the student’s own rollouts, rather than matching only next-token probabilities. It argues this gives lower-variance, denser supervision and reports reaching the teacher on several competition-math benchmarks while training 1.44× faster and using up to 54% less memory. An extension uses learned projector pairs to transfer representations across different model architectures and tokenizers.

    Representation-level on-policy distillation, especially across mismatched tokenizers, is a substantive and potentially reusable direction, but the abstract provides limited evidence beyond a narrow set of math benchmarks and needs scrutiny of comparisons and scaling.

  40. maybe AI / ML ▲ 8 score 4.5

    Breaking the Bubble: Asynchronous Pipeline Parallel Training with Bounded Weight Inconsistency

    Itay Elam, Eliron Rahimi, Avi Mendelson et al.

    The paper proposes PACI, an asynchronous pipeline-parallel training schedule that avoids pipeline bubbles while limiting how many optimizer updates separate a micro-batch’s forward and backward passes. It uses local gradient accumulation to slow parameter-version changes, without weight stashing, extra parameter copies, prediction, or global synchronization; experiments on GPT-style pretraining report comparable perplexity to synchronous training and up to 1.69× better time-to-accuracy than a flush-based baseline.

    Avoiding bubbles without extra model copies while preserving training quality could be a meaningful systems advance, but the abstract provides limited scale, workload, and comparative detail, so the sizable speedup is not yet enough for a strong recommendation.

  41. maybe AI / ML picked▲ 11 score 4.5

    Text-to-Image Models Need Less from Text Encoders Than You Think

    Nurit Spingarn, Noa Cohen, Tamar Rott Shaham et al.

    The authors replace contextual text embeddings in text-to-image diffusion transformers with embeddings that preserve only individual word meanings, subword merging, and word order. They report that this simpler representation produces image quality and prompt fidelity comparable to full text embeddings, suggesting that the image generator—not the text encoder—may handle much of the prompt’s compositional interpretation.

    This is a potentially important and counterintuitive finding about where linguistic composition happens in text-to-image systems, but the abstract provides no quantitative results or evidence about how broadly it holds across models and difficult prompts.

  42. maybe AI / ML ▲ 6 score 4.5

    Neural Networks Provably Learn Spectral Representations for Group Composition

    Jianliang He, Leda Wang, Fengzhuo Zhang et al.

    This paper gives a theoretical account of how a two-layer network trained on finite-group composition develops Fourier/representation-theoretic features. It proves that neurons converge to individual irreducible representations, while cross-layer coefficients align in a low-rank way; for Abelian groups, it further characterizes diversification, phase alignment, and exponential convergence to a majority-vote-like solution.

    The representation-theoretic description of feature learning and the claimed low-rank alignment are a genuinely nontrivial theoretical framing, but the results are confined to an idealized two-layer group-composition setting and their broader relevance to practical networks is unclear.

  43. maybe AI / ML ▲ 30 score 4.5

    World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning

    Yucheng Zhou, Wei Tao, Yiwen Guo et al.

    This paper trains a multimodal language model to decide when to use a visual world-model rollout, assess whether that rollout is reliable, and combine it with abstract reasoning. Its privileged-future self-distillation method uses ground-truth future videos during training but not deployment, and reports roughly 11% gains on two human-verified visual prediction benchmarks with better tolerance to noisy rollouts.

    The explicit controller for invoking and verifying visual simulation, together with privileged-future distillation, is a meaningful approach to combining concrete and abstract reasoning, but the abstract does not establish broad real-world generality or a major capability jump.

  44. maybe AI / ML ▲ 16 score 4.5

    MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?

    Xinyu Che, Junqi Xiong, Yunfei Ge et al.

    The paper studies how to turn noisy, multimodal web guides into executable, editable skills for vision-language agents. Its framework structures the guides, uses the resulting skills to condition a fixed agent, and revises them from trajectory-level failure analysis; across GUI control, gameplay, and card-play tasks, it reports gains of 12.8–25.3 percentage points over vanilla agents across six VLMs. It also finds that raw-guide prompting can hurt, while structured skills and iterative revision are both needed.

    The guide-to-skill formulation and closed-loop skill revision are a meaningful direction with reasonably broad reported gains, but the abstract does not establish whether the improvements reflect a durable agent capability rather than a benchmark- and prompting-specific framework.

  45. maybe AI / ML ▲ 40 score 4.4

    MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism

    Cong Chen, Guo Gan, Kaixiang Ji et al.

    MemDreamer treats hours-long video understanding as iterative exploration rather than feeding the entire video to a VLM. It builds a three-level graph memory of streamed video content, then lets a reasoning model retrieve relevant nodes and relations through an observe–reason–act loop; the authors report a 12.5-point accuracy gain while using only 2% of the full-context input across four benchmarks. The graph-memory and agentic-retrieval combination is useful, though it largely extends established long-context memory and tool-use patterns.

    The reported large accuracy improvement at a tiny fraction of the context is potentially important for long-video systems, but the abstract lacks enough baseline, ablation, and benchmark detail to establish that the gains are more than a well-engineered combination of familiar memory and agentic-retrieval techniques.

  46. maybe Robotics ▲ 4 score 4.4

    Robotic Policy Adaptation via Weight-Space Meta-Learning

    Christian Bianchi, Siamak Yousefi, Alessio Sampieri et al.

    WIZARD uses a meta-learned hypernetwork to generate LoRA adapter weights for a frozen vision-language-action policy from just a language instruction and a short demonstration video. It avoids task-specific action labels and test-time fine-tuning, and reports up to roughly 2× gains on unseen task collections and 14× on unseen tasks in LIBERO, with consistent improvements in a real Franka Panda experiment.

    Generating task-specific policy adapters directly from demonstrations is a meaningful alternative to costly VLA fine-tuning, but the abstract does not provide absolute performance, task counts, or enough detail to establish that the large relative gains generalize beyond LIBERO.

  47. maybe AI / ML ▲ 21 score 4.4

    Decentralized Instruction Tuning: Conflict-Aware Splitting and Weight Merging

    Minsik Choi, Geewook Kim

    The paper proposes MERIT, which partitions heterogeneous instruction-tuning data along principal directions of gradient conflict, trains each partition independently, and merges the resulting models once using token-weighted averaging. A local quadratic analysis argues that this reduces variance and filters undesirable parameter directions; experiments report a gain from 54.3 to 57.0 averaged across eight benchmarks on Qwen2.5-VL-3B, with comparable results on a larger 7B mixture and text-only FLAN.

    The conflict-aware splitting plus one-shot weight merging is a reasonably distinctive approach to reducing synchronization and gradient interference, but the reported capability gain is modest and the abstract provides limited detail about breadth, ablations, and comparison strength.

  48. maybe AI / ML ▲ 7 score 4.4

    Towards One-to-Many Temporal Grounding

    Qi Xu, Yue Tan, Shihao Chen et al.

    The paper introduces one-to-many temporal grounding, where a text query must retrieve multiple disjoint video segments rather than a single interval. It provides a 56k-example dataset, new metrics for counting and temporal coverage, and reward functions that train models to produce more complete, precise segment sets; the proposed model reportedly improves Effective Temporal F1 by about 16 points over strong MLLM baselines.

    This identifies a meaningful capability gap in video-language models—recognizing that an event can recur—and reports a substantial improvement, though the contribution is partly a benchmark, dataset, and task-specific training pipeline and the abstract gives limited evidence beyond one headline metric.

  49. maybe AI / ML ▲ 56 score 4.4

    SoCRATES: Towards Reliable Automated Evaluation of Proactive LLM Mediation across Domains and Socio-cognitive Variations

    Taewon Yun, Hyeonseong Park, Jeonghwan Choi et al.

    SoCRATES is a benchmark for testing LLMs as real-time mediators across eight conflict domains and five variations in participants’ social and emotional conditions. It uses topic-localized scoring, which reportedly aligns with human experts at 0.82, and finds that even frontier models close only about one-third of the gap between unmediated and consensual outcomes, with large weaknesses in adapting to socio-cognitive variation.

    The main interesting result is that current frontier models remain poor at socially adaptive mediation despite strong overall language ability, though this is primarily a benchmark and the abstract provides limited detail about scenario validity and evaluation robustness.

  50. maybe AI / ML score 4.4

    Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning

    Chuang Ma, Qianying Liu, Tomoyuki Obuchi et al.

    This paper identifies a spatial lexical bias in multimodal language models: adding an incorrect spatial relation to the answer choices can cause models to switch from a correct binary answer to the new distractor. Across nine open-weight MLLMs, mechanistic analyses suggest the correct visual relation remains represented, while language-side channels and neurons drive the erroneous choice; a small LLM-only DPO update on synthetic examples substantially reduces the problem, including on several broader datasets.

    The language-side diagnosis of a failure usually attributed to visual grounding, together with large cross-dataset gains from a tiny targeted intervention, is a genuinely non-obvious and potentially useful result, though the abstract does not establish how broadly it generalizes beyond spatial QA.

  51. maybe AI / ML picked score 4.4

    Speculative Sampling For Faster Molecular Dynamics

    Arthur Kosmala, Stephan Günnemann, Meng Gao et al.

    The paper adapts speculative sampling to molecular dynamics: a fast draft force model proposes multiple Langevin integration steps, which are then checked in parallel using a slower target model and corrected with a transport map. The authors provide theory for second-order Langevin dynamics and report 3–9× speedups across systems and draft/target model pairs while preserving the target model’s trajectory distribution.

    This is a genuinely interesting efficiency direction because it attacks MD’s serial bottleneck with an apparently exact speculative procedure, but the abstract does not establish how broadly the speedups hold or how costly the verification and correction steps are.

  52. maybe AI / ML picked score 4.4

    LEAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks

    Po-Nien Kung, Linfeng Song, Dawsen Hwang et al.

    LEAP is an agentic system for producing Lean proofs by combining informal mathematical planning, decomposition, iterative self-refinement, and compiler feedback. It reportedly raises one-shot solve rates on the new Lean-IMO-Bench from under 10% to 70%, solves all 12 problems from the 2025 Putnam competition, and formalizes a subproblem related to Knuth’s Hamiltonian decomposition conjecture.

    The reported jump in formal-proof success and the claimed verified formalization of a research-level combinatorics result are unusually substantial, but the core agentic ingredients are familiar and the abstract does not establish enough detail about benchmark construction, contamination, or independent verification for a strong verdict.

  53. maybe Robotics picked score 4.4

    What Are We Actually Benchmarking in Robot Manipulation?

    Tianchong Jiang, Xiangshan Tan, Samuel Wheeler et al.

    This paper argues that common robot-manipulation benchmark scores are poor proxies for general capability because of shortcut solutions, weak statistical testing, benchmark overfitting, and dependence on the data source. Auditing five benchmarks, the authors find that a small language-free model nearly matches reported LIBERO state of the art, gains on LIBERO are often not statistically significant, and policies trained on CALVIN degrade under modest within-range pose randomization. They provide four reusable diagnostics and reference implementations for evaluating benchmark validity.

    The concrete audits challenge widely used manipulation benchmarks and offer actionable diagnostics, though the abstract does not establish that the conclusions generalize beyond the tested benchmarks or that the proposed tests fully measure real-world manipulation ability.

  54. maybe AI / ML score 4.4

    LDARNet: DNA Adaptive Representation Network with Learnable Tokenization for Genomic Modeling

    Daria Ledneva, Denis Kuznetsov

    LDARNet introduces a genomic foundation model that learns where to place token boundaries rather than using fixed nucleotides, k-mers, or BPE tokens. Its hierarchical architecture combines dynamic chunking with bidirectional state-space and local-attention layers; in experiments it wins most compact-model comparisons and substantially improves histone-modification prediction at matched compute, while learned boundaries align with promoter and splice motifs without supervision.

    The combination of unsupervised adaptive tokenization and genomic modeling is a genuinely interesting direction, and the matched-compute gains and motif alignment are non-obvious, but the abstract alone does not establish broad robustness beyond the reported benchmark suites.

  55. maybe AI / ML picked score 4.4

    The Right Measure for Physics-Constrained Generation: A Co-Area Correction for Posterior-Consistent PDE Inverse Problems

    Jian Xu, Yanning Wu, Delu Zeng et al.

    The paper argues that projecting a generative model onto an exact PDE constraint does not generally produce the correct Bayesian posterior, because conditioning on a measure-zero manifold requires a co-area (Fixman) Jacobian correction. It introduces CoCoS, a sampler incorporating this correction, and reports substantially lower posterior error than projection, guidance, and naive reweighting on controlled problems, matching an independent ground-truth sampler to sampling noise.

    This identifies a potentially fundamental and overlooked measure-theoretic bias in physics-constrained generative inference, with a principled correction and quantitative validation, but the evidence described is still limited to controlled PDE problems rather than broad real-world inverse tasks.

  56. maybe AI / ML picked score 4.4

    Beyond Text Following: Repairable Arbitration Reversals in Audio-Language Models

    Yichen Gao, Yiqun Zhang, Zijing Wang et al.

    This paper argues that audio-language models often extract the correct audio evidence but let conflicting text win during answer arbitration. Using same-audio counterfactuals and activation patching, the authors find preference reversals in 64.1% of conflict cases and localize them to answer-position computation. They introduce a training-free decoding method, GACL, which improves audio faithfulness over contrastive decoding and also transfers to vision-text conflicts.

    The combination of a useful mechanistic diagnosis and a training-free correction with reported cross-modal transfer is genuinely interesting, but the abstract lacks absolute performance numbers and broader evidence needed for a strong recommendation.

  57. maybe AI / ML picked score 4.4

    Dominant-Layer ZO: A Single Layer Dominates Zeroth-Order Fine-Tuning of LLMs

    Wanhao Yu, Ziyan Wang, Zheng Wang et al.

    The paper finds that zeroth-order fine-tuning of LLMs can often be concentrated in one model-specific decoding layer: updating that layer alone matches or outperforms updating the full model. The layer can reportedly be identified before training from activation outliers, and the authors attribute its effect to high perturbation sensitivity plus its early position in the residual stream. On LLaMA2-7B and Qwen3-8B across nine tasks, this gives up to a 4.52× training speedup over full-model ZO baselines.

    The combination of a striking single-layer concentration effect, inference-only layer selection, and substantial zeroth-order fine-tuning speedups is genuinely interesting, though the evidence is limited to two model families and the abstract does not establish how broadly the finding generalizes.

  58. maybe Robotics score 4.4

    Preserving Full 6-DOF Actuation Under Abrupt Total Rotor Failures: Passive Fault-Tolerant Flight Control Using a Biaxial-Tilt Hexacopter

    Yipeng Yang, Yiqiao Tang, Hao Zhang et al.

    The paper studies a hexacopter with independently tilting rotors that can retain full 6-DOF control after certain abrupt, previously unknown rotor failures. It introduces a wrench-space metric that accounts for transient jumps and two passive compensation schemes—one at the controller level and one at control allocation—and reports simulation and flight tests including multi-rotor failures, windy tracking, and contact-based aerial writing. This is a more capable fault-tolerant flight architecture than conventional coplanar or uniaxial-tilt multirotors, though the results apply only to failure cases where the remaining system is still fully actuated.

    The combination of biaxial tilting, no fault detection or mode switching, and demonstrated 6-DOF flight after multiple rotor failures is a substantive real-world capability, but the scope is limited to favorable failure configurations and the abstract gives no quantitative performance margins.

  59. maybe AI / ML picked score 4.4

    The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models

    Kuan-Yen Chen, Fang-Yi Su, Shih-Yen Lin et al.

    The paper tests whether LLMs fail to correct their own mistakes because of a true self-monitoring limitation or because the mistake is labeled as the model’s own thought. Keeping the erroneous text unchanged, the authors present it under different roles such as thought, user, tool output, or memory; external-role relabeling raises explicit correction rates by 23–93 percentage points across most of 12 model-domain settings. The result suggests that chat-template role conditioning, rather than only reasoning ability, substantially controls apparent self-correction.

    The large and consistent effect from changing only the message role is a non-obvious finding with implications for how self-correction and instruction tuning are evaluated, though the abstract does not establish how robust the effect is beyond the tested templates and tasks.

  60. maybe AI / ML score 4.4

    Optimally taming biases in black-box models for efficient semiparametric estimation

    Yihong Gu, Qishuo Yin, Tianxi Cai et al.

    This paper argues that the usual double-machine-learning error bound is not optimal when a nuisance regression, such as E[T|X], cannot be consistently learned. It proposes an estimator whose target-parameter error scales as n^-1/2 plus nuisance approximation error plus the squared stochastic error, proves a matching lower bound, and extends the idea beyond the partial linear model to functionals such as average treatment effects. The practical implication is to undersmooth nuisance learners rather than use the usual bias–variance balance.

    The claimed elimination of first-order nuisance stochastic error without extra assumptions, together with an optimality lower bound, would materially revise how black-box nuisance estimation is analyzed and tuned in semiparametric inference, though the abstract gives no empirical validation and the scope beyond the core model is not yet clear.

  61. maybe AI / ML score 4.4

    How abundant are good interpolators?

    August Y. Chen, Ahmed El Alaoui

    This paper studies the distribution of test performance among all unit-norm linear classifiers that interpolate a dataset, rather than analyzing only a particular training algorithm. For Gaussian-mixture and logistic models in a proportional overparameterized regime, it proves a large-deviation principle showing that almost all interpolating classifiers have performance near a deterministic typical value; gradient descent and a linear program can nevertheless find exponentially rare interpolators that generalize substantially better. The result gives a rigorous geometric explanation for why benign overfitting can depend strongly on the optimization procedure.

    The paper makes a non-obvious and theoretically substantial distinction between the typical interpolator and algorithmically selected interpolators, with rigorous high-dimensional results, though its scope is limited to specific linear models and low sample-to-dimension ratios.

  62. maybe AI / ML score 4.4

    The Sharp Phase Transition of Tyler's M-Estimator for Robust Subspace Recovery

    Gilad Lerman, Teng Zhang

    This paper analyzes Tyler’s M-estimator for robust subspace recovery exactly at the critical dimension-scaled signal-to-noise ratio (DS-SNR) of 1. It proves that, under a stability condition weaker than prior general-position assumptions, the estimator converges to the true subspace even at the boundary, establishing a sharp transition between failure and successful recovery.

    The boundary-case convergence result links a practical estimator to the known computational-hardness threshold and weakens the required assumptions, but it remains a specialized theoretical advance rather than a broad new capability.

  63. maybe AI / ML score 4.4

    RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation

    Yuan Si, Simeng Han, Daming Li et al.

    RENDER isolates the effect of how stored conversational information is presented to an LLM, while keeping the underlying history and answer unchanged. Across 500 LongMemEval questions and nine models, natural-language memory entries substantially outperform raw or typed-record formats—sometimes by tens of percentage points—and the effect persists with retrieval noise and on HotpotQA. The main contribution is showing that reader-facing serialization is a major, often uncontrolled variable in memory and RAG evaluation.

    The large and representation-dependent performance gaps expose a non-obvious confound in memory/RAG research, though this is primarily an evaluation finding rather than a new capability or method, and model-specific significance is mixed.

  64. maybe AI / ML score 4.4

    Less Context, More Accuracy: A Bi-Temporal Memory Engine for LLM Agents Where a Lean Retrieved Context Beats the Full History

    Liuyin Wang

    Engram is a long-term memory system for LLM agents that stores raw episodes, asynchronously extracts time-aware facts, tracks contradictions and provenance, and combines graph, lexical, dense, and recency retrieval. On all 500 questions in LongMemEval_S, it answers from about 9.6k retrieved tokens and reaches 83.6% versus 73.2% for feeding the full 79k-token history, with a reproducible harness and raw logs. The main finding is that a carefully constructed, compact retrieval context can be both cheaper and more accurate than full-history prompting.

    The substantial, statistically supported improvement over full-context prompting at roughly one-eighth the token cost is genuinely notable, but the system combines several established memory and retrieval mechanisms and the evidence is concentrated on one benchmark, so it falls short of a strong verdict.

  65. maybe AI / ML score 4.4

    Unlocking Latent Value: Taxonomy-Guided Recovery of High-Performing Data from Low-Tier Web Corpora

    Neeraj Varshney, Sanket Lokegaonkar, Nasser Zalmout et al.

    The paper argues that web-data quality is not one-dimensional: documents discarded by aggregate quality scores can still be valuable for reasoning, coding, knowledge, timeliness, or cultural specificity. It adds taxonomy dimensions, distills large-model annotations into smaller classifiers, and uses a two-pass search over filter combinations; filtered mid- and low-tier data reportedly beats unfiltered higher-tier data, with gains up to 22.3% on reasoning and 19.5% on coding.

    The substantial result that carefully selected low-tier web data can outperform production-quality tiers is both non-obvious and practically important, but the abstract does not establish how broadly the gains transfer across model scales, datasets, and evaluation setups.

  66. maybe AI / ML score 4.4

    PACE: Anytime-Valid Acceptance Tests for Self-Evolving Agents

    Zayx Shawn

    PACE treats accepting self-generated agent changes as a sequential hypothesis-testing problem rather than simply keeping any change that improves a noisy development score. It uses paired evaluations and an anytime-valid betting test to control false commits under repeated, optional evaluation, and on several Qwen2.5 agents and reasoning benchmarks it sharply reduces harmful or spurious edits while retaining genuine improvements and lowering evaluation cost.

    The acceptor-focused framing and anytime-valid statistical gate address a real failure mode in iterative self-evolution, with unusually concrete results, though the core testing machinery is adapted from established sequential statistics and the evidence is limited to prompt-level evolution on a few benchmarks.

  67. maybe AI / ML score 4.4

    How Deep Are Deep GPs, Really? A Sharp Threshold and a Non-Gaussian Limit for Compositional GPs

    Mark Kozdoba, Shie Mannor

    This paper analyzes what happens when Gaussian-process layers are composed to very great depth. It proves a sharp RBF-kernel bandwidth threshold scaling as the square root of input dimension: above it, the prior collapses to constant functions, while below it, the composition converges to a non-degenerate, non-Gaussian, coordinate-dependent distribution. Experiments verify the threshold and show multimodal limiting behavior in the narrow non-degenerate regime.

    The sharp phase transition and proof of non-Gaussian deep limits provide a genuinely useful theoretical picture of when deep GP priors remain expressive, but the result is focused on a specific compositional GP setting rather than an immediate capability or practical-model breakthrough.

  68. maybe AI / ML picked score 4.4

    Contemporary AI lacks the imagination to diverge or negate in science

    Honglin Bao, Siyang Wu, Xiao Liu et al.

    The authors collect 25,139 scientist evaluations of LLM-generated research ideas derived from 121,640 papers, comparing reasoning and non-reasoning models across four scientific fields. They report that ordinary LLMs produce highly similar ideas, reasoning models broaden exploration but rarely generate null hypotheses, and automated judges correlate weakly with experts; a reward model trained on human ratings performs substantially better. Linking this to 39 million papers, they argue that post-ChatGPT scientific ideas have become less divergent and that human grounding remains important.

    The unusually large scientist-in-the-loop evaluation and the specific finding that LLMs suppress null-hypothesis thinking and produce a post-ChatGPT contraction of ideas are genuinely interesting, though the macro-level causal interpretation and reward-model comparison need careful scrutiny.

  69. maybe Tech score 4.4

    Enhanced Wide-Angle Steering with Multi-Mode Multi-Port Aperture Antenna Arrays

    Tim Hahn, Dirk Manteuffel

    The paper introduces an antenna array whose aperture elements use multiple modes and ports, together with a beamforming method that exploits those extra degrees of freedom. A fabricated prototype reportedly scans to ±77° from broadside with only 3 dB scan loss in both principal planes and without visible grating lobes, supported by measurements.

    The measured combination of very wide-angle two-dimensional scanning and low scan loss is a substantial antenna-engineering result, though the contribution is specialized and the abstract gives too little detail for a strong recommendation.

  70. maybe AI / ML score 4.4

    Learning to Solve Generative ODEs Beyond the Linear Span

    Sihyeon Kim, Seunghun Lee, Vikas Singh et al.

    The paper argues that learned ODE solvers are fundamentally limited when each update is only a scalar recombination of previously evaluated velocities: they cannot produce state corrections outside that span. SpanLift adds a lightweight learned spatial residual operator on top of a fixed solver, trained by endpoint matching without additional backbone evaluations. It reports large 3-NFE improvements, including CIFAR-10 FID 8.16→5.69 and ImageNet FID 17.37→11.83, across diffusion, flow matching, and precipitation forecasting.

    The explicit span-bottleneck diagnosis and no-extra-NFE residual correction are a genuinely interesting direction, with unusually large reported few-step gains across several settings, but the abstract lacks enough comparative and scaling detail to justify a strong verdict.

  71. maybe AI / ML score 4.4

    APEX4: Efficient Pure W4A4 LLM Inference via Intra-SM Compute Rebalancing

    Hong Guo, Nianhui Guo, Weixing Wang et al.

    APEX4 studies why pure INT4 weights/activations (W4A4) quantization is fast on some GPUs but slower on others, attributing the difference mainly to the Tensor Core–CUDA Core throughput ratio within an SM. It uses that hardware signal to choose quantization granularity and rebalance kernel work, achieving near-FP16 perplexity and up to 1.66–2.09× end-to-end speedups in vLLM, while recovering a smaller 1.20–1.40× gain on A100.

    The hardware-dependent explanation and ρ-aware pure-W4A4 kernel design are more than a routine quantization tweak, with broad GPU measurements and substantial real inference speedups, though the contribution remains a specialized systems optimization rather than a new general ML capability.

  72. maybe Neuroscience picked score 4.4

    Prenatal assembly of functional cortical circuits

    Morassut, I., Panzeri, A., Fievre, S. et al.

    The authors compare precocial Acomys mice, which develop in utero for 39 days, with altricial Mus mice, which gestate for 19 days, using anatomy, birth dating, electrophysiology, transcriptomics, and behavior. They find that Acomys follows the same overall cortical developmental program but completes major neuronal, circuit, molecular, and sensorimotor milestones before birth, suggesting that birth itself is not required for early cortical maturation.

    This is a strong comparative test of whether postnatal experience drives cortical maturation, with convergent evidence for the non-obvious conclusion that developmental programs can unfold prenatally, but the abstract does not establish a direct mechanism or broader cross-species generality.

  73. maybe Neuroscience score 4.4

    (2R,6R)-Hydroxynorketamine elicits rapid antidepressant effects by promoting astrocytic μ-δ opioid receptor heterodimerization

    Liang, Y., Wang, L., Li, Y. et al.

    The study identifies astrocytic μ-δ opioid receptor heterodimers as an upstream target of (2R,6R)-hydroxynorketamine (HNK), a ketamine metabolite with rapid antidepressant-like effects but no NMDAR blockade. Using pharmacology, genetics, super-resolution imaging, simulations, and receptor mutations, the authors argue that HNK binds the μ-receptor and restores stress-reduced heterodimerization, which is required for synaptic and behavioral effects in mice.

    This is a specific and potentially important mechanistic link between a ketamine metabolite, astrocytic opioid-receptor heterodimers, and rapid antidepressant action, supported by several converging methods, but the evidence is still preclinical and the therapeutic significance remains unestablished.

  74. maybe Neuroscience picked score 4.4

    Cortical folding patterns are encoded in the geometry of the unfolded neocortex.

    Toro, R., Heuer, K., Aflak, N. et al.

    The authors track ferret cortical surface geometry from birth, before visible folding, through adulthood. They report that newborn curvature predicts where sulci and gyri will form, their orientation, and mature curvature; gene-expression associations and mechanical simulations suggest that early geometry acts as an organizing signal alongside molecular patterning.

    This presents a potentially important shift from a primarily molecular account of cortical folding to one in which pre-existing geometry and mechanics provide positional information, supported by developmental imaging and simulations, though the evidence is limited to ferret development and abstract-level details.

  75. maybe AI / ML ▲ 53 score 4.4

    Benchmarking Visual State Tracking in Multimodal Video Understanding

    Sihyun Yu, Nanye Ma, Pinzhi Huang et al.

    The authors introduce VSTAT, a benchmark of 834 synthetic and real-world video clips with 1,500 questions that require tracking visual states and events across time rather than inspecting a single frame. State-of-the-art multimodal models perform far below humans and only modestly better than answer-prior baselines; analysis suggests their textual reasoning can be correct, but they fail to perceive and track the necessary visual events. Preliminary tests indicate that video and coding agents do not substantially fix this weakness.

    This is primarily a benchmark paper, but its non-obvious diagnosis—that current MLLMs’ temporal reasoning failures are often perceptual tracking failures, and that agentic scaffolding does not readily solve them—makes it worth a closer look.

  76. maybe AI / ML ▲ 31 score 4.4

    OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs

    Yifei Li, Pengyiang Liu, Yuhang Zang et al.

    OVO-S-Bench evaluates multimodal models on spatial reasoning from streaming egocentric video, where models can access only the video prefix available before each query. Across 38 models, even Gemini-3.1-Pro remains far below human performance, with allocentric mapping hardest; spatial fine-tuning and ungrounded chain-of-thought can actually reduce performance relative to the base models. The main contribution is a carefully structured benchmark plus analyses exposing failures that ordinary full-video evaluations miss.

    Although primarily a benchmark paper, it reports non-obvious results—especially that spatial fine-tuning and chain-of-thought can worsen streaming spatial reasoning—and uses broad model comparisons with human-prefix-protocol baselines.

  77. maybe AI / ML ▲ 50 score 4.4

    ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?

    Woojung Song, Nalim Kim, Sangjun Song et al.

    The paper introduces ArcANE, a benchmark for testing whether role-playing language models follow a character’s changing psychological trajectory rather than merely recalling facts or maintaining a fixed persona. Across 17 novels, 80 characters, six models, and six context strategies, explicitly providing a segmented character arc performed best, especially on new scenarios absent from the source text; fine-tuning on the benchmark data increased this advantage.

    The focus on extrapolating evolving character psychology beyond retrieved source text is a useful and somewhat non-obvious evaluation direction, but the abstract gives no effect sizes and the benchmark’s automatically constructed labels and fine-tuning setup need scrutiny.

  78. maybe Robotics ▲ 10 score 4.4

    PlatonicNav: Unveiling Semantic Correspondence in Navigation with Platonic Topological Maps

    Junlin Long, Zeyu Zhang, Xu Deng et al.

    PlatonicNav argues that vision-only representations may already contain enough semantic structure to support language-conditioned navigation, without paired image-text training. It builds an object-centric topological map from a self-supervised visual encoder and matches language goals to the map using a training-free procedure, reporting results across object-goal and vision-language navigation benchmarks plus a Unitree Go2 deployment.

    The potentially important result is zero-shot language grounding from a purely vision-built map, but the abstract gives no quantitative gains or ablations, so the strength and generality of the claimed shared semantic manifold are unclear.

  79. maybe AI / ML ▲ 3 score 4.4

    Large Language Models Are Overconfident in Their Own Responses

    Mario Sanz-Guerrero, Manuel Mager, Katharina von der Wense

    The paper finds that instruction-tuned chat models are not only less calibrated overall, but are especially confident when evaluating answers they generated themselves. Across six open-weight models, three benchmarks, and three confidence methods, models gave their own answers up to 26% more confidence than identical answers attributed to users. Reformatting the model’s answer as user-provided text at inference time substantially improves calibration without retraining.

    The ownership-bias finding is a non-obvious explanation for conversational-model miscalibration and is supported across multiple models and evaluation settings, though the proposed fix is a narrow inference-time prompt-formatting trick rather than a major capability advance.

  80. maybe AI / ML ▲ 3 score 4.4

    Reinforcement Learning from Rich Feedback with Distributional DAgger

    Rishabh Agrawal, Jacob Fein-Ashley, Paria Rashidinejad

    The paper proposes DistIL, a distributional version of DAgger for training reasoning models from rich, step-level feedback rather than only binary final-answer rewards. It uses forward cross-entropy against an expert action distribution, propagating later expert–student disagreement back to earlier decisions; the authors prove monotonic policy-improvement and regret guarantees and report gains over RLVR and self-distillation on reasoning, coding, and mathematics tasks.

    The combination of distributional DAgger, black-box rich feedback, and a theoretically justified alternative to reverse-KL self-distillation is a meaningful direction, but the abstract gives no quantitative results or detail on the breadth of the empirical validation.

  81. maybe AI / ML ▲ 5 score 4.4

    SEAOTTER: Sensor Embedded Autoencoding with One-Time Transcode for Efficient Reconstruction

    Dan Jacobellis, Neeraja J. Yadwadkar

    SEAOTTER is a learned image-compression pipeline for robotics sensors that converts a compact neural representation into a standard JPEG, preserving compatibility with existing infrastructure. It jointly learns the JPEG color and quantization transforms for either general reconstruction or downstream perception, reporting 200:1 compression with 7× faster encoding, 3.5× faster decoding, and 8 percentage points higher ImageNet top-1 accuracy than AVIF.

    The combination of asymmetric learned compression with a one-time, learnable transcode into standard JPEG is a non-obvious systems idea with potentially broad deployment value, but the abstract gives limited evidence beyond one benchmark and unclear accuracy baselines.

  82. maybe AI / ML ▲ 5 score 4.4

    Towards Retrieving Interaction Spaces for Agentic Search

    Shengyao Zhuang, Yuansheng Ni, Hengxin Fun et al.

    This paper reframes retrieval for search agents as constructing a bounded, tool-enabled “interaction space,” rather than merely returning a few documents for the model to read. Its RISE prototype uses BM25 to select documents and preprocesses them for shell-style navigation; on BrowseComp-Plus it matches a pure-shell baseline at about one-quarter the cost, and remains effective at 1M documents where unbounded interaction becomes slow and unreliable.

    The interaction-space framing and demonstrated cost/scaling advantage are genuinely interesting, though the method is currently a BM25-based proof of concept evaluated mainly on one benchmark.

  83. maybe AI / ML ▲ 15 score 4.4

    Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators

    Chenming Zhu, Jingli Lin, Yilin Long et al.

    The paper equips a vision-language model with a learned world simulator that generates alternative views from images and natural-language camera movements, allowing the model to gather imagined visual evidence during reasoning. Its RL-trained agent learns when and how to invoke the simulator; on two spatial-reasoning benchmarks, the combined system improves substantially over the underlying VLMs, though the evidence is limited to these reported evaluations.

    The combination of action-conditioned visual imagination, view-consistent world simulation, and learned tool-use policy is a meaningful direction for spatial reasoning, but the abstract does not establish broad capability gains beyond benchmark improvements.

  84. maybe AI / ML ▲ 25 score 4.4

    Reinforcement Learning Elicits Contextual Learning of Unseen Language Translation

    Hanxu Hu, Zdeněk Šnajdr, Pinzhen Chen et al.

    The paper trains LLMs with reinforcement learning to translate extremely low-resource or unseen languages from rich linguistic information supplied in context. Unlike continued training or ordinary in-context prompting, the goal is to learn a general skill for extracting and applying unfamiliar grammars; the authors report better translation on entirely unseen languages using chrF-based rewards.

    Learning a transferable meta-skill for language acquisition through outcome-based RL is a meaningful idea, but the abstract gives no quantitative results or evidence about breadth, scaling, or robustness, so it is not strong enough to prioritize highly.

  85. maybe AI / ML ▲ 24 score 4.3

    LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing

    Jianzong Wu, Hao Lian, Jiongfan Yang et al.

    LoomVideo is a 5B-parameter diffusion-transformer system that handles multimodal video generation and editing in one model. Its main efficiency idea is to condition editing by scaling and adding the source-video latent to the target latent, rather than concatenating source tokens, reportedly delivering at least 5.41× faster inference while retaining complex editing ability; it also uses an MLLM encoder and multi-layer feature injection for multimodal inputs.

    The zero-overhead latent conditioning could be a useful and broadly applicable way to reduce the cost of unified video editing, but the abstract provides few concrete benchmark results beyond generic state-of-the-art claims and the reported speedup needs careful validation.

  86. maybe AI / ML ▲ 1 score 4.3

    Cosine Misleads: Auxiliary Losses Reshape Vision Language Models, Not Their Latents

    XiuYu Zhang, Junfeng Fang, Zhenkai Liang

    The paper tests whether supervised latent tokens in vision-language models actually carry the visual reasoning signal they are trained to represent. Across five variants, it finds that better cosine alignment is strongly associated with worse accuracy, while probing and corruption experiments suggest the model often bypasses the latent and instead changes shared language-model parameters; the latent affects accuracy by at most four points when corrupted.

    The strong, counterintuitive finding that the standard latent-alignment metric is anti-correlated with performance—and that the optimized latent may be largely bypassed—is worth checking, but the evidence is limited to five variants and abstract-level claims.

  87. maybe AI / ML ▲ 22 score 4.3

    UnpredictaBench: A Benchmark for Evaluating Distributional Randomness in LLMs

    Amirhossein Abaskohi, Amirhossein Dabiriaghdam, Liang Luo et al.

    UnpredictaBench evaluates whether LLMs can generate samples matching a specified probability distribution, rather than merely producing varied answers. Across 448 problems spanning standard distributions, stochastic programs, and natural-language random processes, models perform poorly: KS@100 scores range from near 0 to above 20%, and no tested model exceeds 40%; chain-of-thought helps somewhat but does not solve the problem.

    This is a useful and relatively neglected capability evaluation with evidence that LLMs’ apparent diversity does not imply calibrated stochastic simulation, though it remains primarily a benchmark and the abstract does not establish a major methodological breakthrough.

  88. maybe AI / ML ▲ 4 score 4.3

    AsyncWebRL: Efficient Asynchronous Reinforcement Learning for Multi-Step Visual Web Agents

    Hao Bai, Rui Yang, Chenlu Ye et al.

    AsyncWebRL combines an asynchronous training pipeline with a change to multi-step GRPO normalization for visual web agents. The system reportedly improves end-to-end training throughput by up to 2.9×, while replacing length-dependent trajectory weighting shortens failed trajectories without reducing overall success; together, the method improves WebGym OOD performance from 42.9% by 5.8% relative, with larger gains on harder splits.

    The potentially reusable insight is that length-normalizing trajectories can systematically weaken learning from verbose failures, but the evidence is mainly on one web-agent benchmark and the claimed gains need validation beyond the authors' pipeline.

  89. maybe AI / ML ▲ 2 score 4.3

    ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMs

    Ashutosh Hathidara, Sai Shruthi Sistla, Sebastian Schreiber et al.

    ToolSense tests whether LLMs genuinely understand tools rather than merely exploiting fully specified queries and constrained decoding. On ToolBench’s roughly 47,000 tools, several parametric-retrieval configurations lost 50–64 percentage points on more realistic ambiguous queries, and some performed near chance on factual probes despite strong standard retrieval scores.

    The reported gap between benchmark retrieval success and actual tool knowledge is a useful, non-obvious diagnostic finding, but this is primarily a benchmark/auditing framework and the validity of automatically generated probes limits how strongly to interpret it.

  90. maybe AI / ML ▲ 2 score 4.3

    Data-Efficient Autoregressive-to-Diffusion Language Models via On-Policy Distillation

    Xingyu Su, Jacob Helwig, Shubham Parashar et al.

    The paper converts a pretrained autoregressive language model into a diffusion language model without full diffusion pretraining. Its student uses bidirectional attention and learns from the original frozen model on trajectories generated by the student itself, addressing both knowledge loss and the mismatch between randomly masked training data and confidence-based inference; the authors report using 15–7,000× fewer training tokens while retaining strong performance across many tasks.

    On-policy distillation as a practical AR-to-diffusion conversion strategy is a meaningful idea with potentially large training-cost implications, but the abstract gives no concrete baselines, task results, or details supporting the very broad performance claim.

  91. maybe AI / ML ▲ 7 score 4.3

    Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning

    Atoosa Chegini, Soheil Feizi

    The paper uses a large, frozen language model to score fixed-length chunks produced by a smaller model during mathematical reasoning, committing one chunk at a time instead of selecting only among complete answers or training a process reward model. Its contrastive scoring rule improves over majority voting by as much as 28 percentage points and often matches or beats a much larger trained PRM, while producing shorter reasoning traces; it also identifies persistent length bias when scoring variable-length steps.

    The training-free chunk-level and contrastive process-scoring setup is a meaningful alternative to PRM-guided search with broad benchmark evidence, but the demonstrated gains are confined to mathematical reasoning and may depend heavily on the chosen model pair and inference budget.

  92. maybe AI / ML score 4.3

    Don't Let a Few Network Failures Slow the Entire AllReduce

    Peiqing Chen, Jiedong Jiang, Nengneng Yu et al.

    The paper studies AllReduce when one or more servers lose network bandwidth but remain available. It derives a lower bound showing that, if a straggling server retains at least half its bandwidth, the unavoidable slowdown is only O(1/p), then proposes OptCC, a pipelined algorithm that reportedly stays within 2–6% of fault-free ring performance under up to 50% bandwidth loss, versus as much as 57% overhead for existing methods. The main evidence comes from SimAI simulations rather than large-scale hardware experiments.

    The combination of an information-theoretic bound and a fault-tolerant collective that largely avoids putting degraded servers on the critical path is a meaningful systems contribution, but the evidence is simulation-only and the result is specialized to a particular failure regime.

  93. maybe AI / ML score 4.3

    An Algebraic View of the Expressivity of Recurrent Language Models

    Franz Nowak, Ryan Cotterell, Reda Boumasmoud

    This paper gives a unified algebraic framework for determining which formal languages recurrent neural language models can recognize, explaining why prior work obtained conflicting results such as regular-language expressivity versus Turing completeness. It relates expressivity to algebraic structures such as syntactic monoids and shows that arithmetic semantics matter: a diagonal state-space model cannot implement even-modulus counters with floating-point recurrences, but can implement all such counters with unsigned-integer quantization.

    The explicit separation between architecture and arithmetic model, together with an algebraic resolution of conflicting expressivity claims, is a genuinely non-obvious theoretical contribution, though the abstract does not provide enough detail to establish broad practical impact.

  94. maybe AI / ML score 4.3

    OctoT2I: A Self-Evolving Agentic Text-to-Image Router

    Xu Jiang, Bin Chen, Gehui Li et al.

    OctoT2I uses a stateful agent to route prompts among multiple text-to-image tools over several rounds, while building its own capability knowledge base without human annotations. Its self-evolving propose–solve–evaluate–learn loop reportedly reaches 0.96 on GenEval while reducing inference time by 90.3% and energy use by 56.6% versus Flow-GRPO.

    The combination of unsupervised capability discovery, adaptive multi-round routing, and a very large claimed efficiency gain is plausibly important, but the abstract does not provide enough experimental detail to establish that the gains are broad or fairly compared.

  95. maybe AI / ML score 4.3

    Learning Implicit Bias in Generative Spaces for Accelerating Protein Dynamics Emulation

    Kaihui Cheng, Zhiqiang Cai, Wenkai Xiang et al.

    The paper adds a history-dependent, distance-based bias to a frozen generative emulator of protein dynamics, steering sampling away from previously visited structures while periodically projecting samples back onto the learned data manifold. On DynamicPDB-80 it increases diversity by 35%; on 12 zero-shot proteins it reaches the baseline emulator’s coverage up to 37× faster and finds roughly 3× more low-energy states.

    This is a plausible new direction for turning generative protein emulators into enhanced-sampling tools, with substantial reported speedups, but the evidence is limited to a small protein set and one main benchmark, so it falls short of a strong recommendation.

  96. maybe AI / ML score 4.3

    Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving

    Jianru Ding, Ryien Hosseini, Pouya Mahdi Gholami et al.

    This paper argues that agentic workloads should be scheduled at the conversation level rather than one turn at a time. Its ConServe system sends the first-turn prefill to a specialized GPU pool, transfers the KV cache once, and keeps the conversation on one decoder thereafter, avoiding predictions of future decode and tool-call behavior; it reports 51.08% lower p95 time-to-first-effective-token and 7.51% better energy efficiency than a per-turn prediction baseline, with further energy gains on heterogeneous GPUs.

    The conversation-level scheduling abstraction and prediction-free two-phase placement are a non-obvious systems idea with substantial reported latency and energy gains, though the abstract provides limited detail on workload breadth and baseline strength.

  97. maybe AI / ML score 4.3

    Extreme Low-Bit Inference in Reasoning Models: Failure Modes and Targeted Recovery

    Ekaterina Alimaskina, Darya Rudas, Denis Shveykin et al.

    The paper studies why aggressive 2-bit quantization hurts reasoning models beyond ordinary accuracy loss: generation becomes unstable, producing loops, overlong traces, delayed answers, and unfinished reasoning. It proposes giving the quantized model a short FP16 plan and/or detecting loops to commit early or fall back to FP16; on MATH-500, loop rescue raises Qwen3-8B accuracy from 17.2% to 74.2%, while planning plus rescue raises Qwen3-32B from 65.0% to 87.2%.

    The process-level diagnosis that extreme quantization causes pathological reasoning trajectories, together with very large reported recovery gains, is substantially more interesting than a standard quantization tweak, but the evidence is limited to Qwen3 models and a small set of benchmarks, with end-to-end speedup details not quantified in the abstract.

  98. maybe AI / ML score 4.3

    Aligning Data-Driven Predictors with Allocation: A Decision-Focused Approach to Survival Analysis

    Itai Zilberstein, Ioannis Anagnostides, Tuomas Sandholm

    The paper argues that survival models optimized for predictive metrics such as the C-index can perform arbitrarily badly when their rankings are used to allocate scarce resources. It proposes optimizing NDCG instead, proves that this ranking objective gives allocation-utility guarantees, and reports 50–100% NDCG gains on historical US heart-transplant data, with a method for handling right censoring.

    The decision-focused reframing and formal link between survival ranking and allocation are substantial, but the very large claimed life-year gains rely on retrospective transplant data and the abstract does not establish that they would survive causal or deployment-level scrutiny.

  99. maybe Robotics score 4.3

    Intercepting the Future: Latent-Space Predictive World Model for Dynamic VLA Manipulation

    Shahram Najam Syed, Arthur Jakobsson, Haoran Hao et al.

    The paper adds a small predictive world-model wrapper to a frozen 7B OpenVLA, forecasting future VLA feature tokens from optical-flow-derived motion and using them to choose actions before moving objects reach the grasp point. Across dynamic manipulation simulations it reports 79–97% success versus 31–58% for the strongest baseline, and on a physical xArm 7 it handles conveyor, rolling-ball, interception, and catching tasks, including 19/30 projectile catches where baselines achieve 0/30.

    The combination of latent-space prediction, adaptive anticipation, and a frozen VLA addresses a real failure mode and the physical catching results are striking, but the abstract does not establish how broad or robust the gains are beyond a small set of designed tasks.

  100. maybe AI / ML score 4.3

    SimSD: Simple Speculative Decoding in Diffusion Language Models

    Junxia Cui, Haotian Ye, Runchu Tian et al.

    SimSD adapts speculative decoding to diffusion language models, whose bidirectional denoising context normally prevents standard token-level verification. It uses draft-model reference tokens and a specially designed attention mask to make drafted-token logits temporally valid, without retraining the models. On SDAR-family models and four benchmarks, it reports up to 7.46× higher decoding throughput while preserving or improving generation quality.

    This addresses a real architectural obstacle to bringing a major autoregressive inference technique to diffusion LMs, with a potentially large reported speedup, but the abstract does not establish how broadly the gains hold or how they compare with existing diffusion-specific acceleration methods.

  101. maybe AI / ML score 4.3

    EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement

    Hui Li, Yangfan Gao, Junlin Shang et al.

    EntangleCodec is a discrete audio tokenizer designed to represent semantic content and acoustic details in one compact token stream, rather than separate semantic and acoustic streams. It uses caption-aligned representations and a flow-matching decoder, and reportedly matches specialized codecs for reconstruction while improving audio understanding by up to 7.4%; models using it also outperform much larger continuous-representation baselines on several benchmarks. The main novelty is showing that a better unified audio representation can substantially reduce the language-model scale needed for audio understanding and generation.

    The unified semantic-acoustic representation and reported 22× parameter-efficiency advantage are genuinely interesting for audio language models, but the abstract provides limited detail about comparisons, ablations, and whether the large gains generalize beyond the cited benchmarks.

  102. maybe AI / ML score 4.3

    Global Unknown Estimation: A Statistical Framework for Wireless Distributed Learning

    Yicheng Qu, Ali Bereyhi, Ben Liang

    The paper reframes wireless model aggregation as statistical inference rather than direct AirComp computation. Its global unknown estimation (GUE) method models the relationship between local and global parameters and reportedly reduces the required aggregation power by about 15 dB in low-SNR settings, without extra computation.

    The inference-based formulation and large low-SNR power reduction are potentially important, but the abstract provides only numerical validation and too little detail about assumptions, workloads, or robustness to justify a stronger verdict.

  103. maybe AI / ML score 4.3

    Depth from Dual Differential Defocus and Stereo Consensus

    Junjie Luo, Wei Xu, Dylan Chu et al.

    The paper combines a new dual-defocus depth cue with stereo, using agreement among several physically independent depth estimates to select reliable predictions. A prototype with only a 4 mm baseline and 12 mm focal length reportedly produces 900×1800 depth maps with 1 cm mean absolute error from 0.3–1.64 m in a single snapshot, while maintaining a working range comparable to stereo systems with baselines about 10× larger.

    The compact passive depth-sensing design and claimed 10× baseline reduction are genuinely notable, but the abstract provides limited comparative and experimental detail, so the strong hardware and accuracy claims need verification.

  104. maybe AI / ML score 4.3

    Scalable Derivative Gaussian Processes via Exact Gradient Reduction

    Hyunseok Seung, Matthias Katzfuss

    The paper introduces TERA, a Vecchia-style approximation for Gaussian processes with function and gradient observations. It shows that, for stationary kernels, only a small set of directional gradient components—at most m² for m conditioning points—affect the conditional distribution of a target function value, allowing the full derivative GP model to be retained while avoiding cubic dependence on the ambient dimension. The resulting method claims nearly dimension-independent runtime and memory, with orders-of-magnitude speedups and comparable or better predictive accuracy.

    The exact, target-specific elimination of irrelevant gradient directions is a genuinely interesting structural result that could make derivative GPs practical in high dimensions, but the abstract provides limited quantitative evidence and relies on a Vecchia approximation rather than fully scalable exact inference.

  105. maybe AI / ML picked score 4.3

    PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation

    Yichuan Wang, Zhifei Li, Zirui Wang et al.

    PixelRAG retrieves and reads webpages as screenshots rather than converting them into text, using a visual embedding index over 30 million Wikipedia screenshots and a vision-language model for answer generation. The authors report consistent gains over text-based RAG on text-only, multimodal, noisy-web, and agentic QA tasks—up to 18.1%—and show that lowering screenshot resolution can cut token costs by up to 3x without hurting accuracy.

    This is a genuinely nonstandard RAG representation with potentially broad implications for layout- and multimodal-aware retrieval, but the abstract gives limited detail about benchmark sizes, baseline quality, and whether the gains generalize beyond carefully curated screenshot training.

  106. maybe AI / ML score 4.3

    Inducing Reasoning Primitives from Agent Traces

    Zhihan Lei, Jiarui Yan, Joshua Momo et al.

    The paper mines successful ReAct traces, clusters recurring reasoning steps, and turns them into a small library of typed pseudo-tools described in natural language. Reusing these induced primitives substantially improves performance over the trace-generating agent and zero-shot chain-of-thought on five reasoning and planning subtasks, with reported gains of 22–44 percentage points and lower inference cost than AWM.

    The promising idea is to convert recurring agent reasoning patterns into reusable, automatically induced abstractions, but the evidence is limited to benchmark-style tasks and the abstract does not establish robustness beyond those settings.

  107. maybe AI / ML score 4.3

    ACRONYM: Accelerated Approximate Nearest Neighbor Search in Memory for Dynamic Vector Databases

    Md Mizanur Rahaman Nayan, Tianqi Zhang, Flavio Ponzina et al.

    ACRONYM is an algorithm–hardware co-design for approximate nearest-neighbor search in dynamic vector databases. It uses data-distribution-independent binary encoding, CAM-based in-memory Hamming-distance search, approximate top-k selection, and coarse-to-fine refinement so updates can occur continuously without index rebuilding. The authors report over 90% recall at 8 million queries/s on million-scale datasets, using 32 MB and 2.56 μJ per query, with large claimed speedups over CPU HNSW and GPU FAISS-IVF.

    The combination of update-friendly binary indexing and specialized in-memory search could address a real bottleneck in dynamic vector retrieval, but the very large hardware-relative speedup claims need careful scrutiny of recall definitions, workloads, baselines, and implementation assumptions.

  108. maybe AI / ML score 4.3

    GFFMERGE: Efficient Merging of Graph Neural Force Fields and Beyond

    Parth Verma, Parv P. Singh, Vipul Garg et al.

    GFFMERGE merges separately trained graph neural force fields by analytically aligning their message-passing embeddings, avoiding full retraining when adapting to new chemical systems. Across molecular, solid-state, and general graph benchmarks, it reportedly outperforms standard vision/language model-merging methods, reaches near joint-training performance, and provides 5–27× faster adaptation with better initialization for fine-tuning.

    The closed-form embedding-alignment approach and the reported failure of existing merging methods on force-field regression are genuinely non-obvious, but the abstract gives no quantitative errors or details sufficient to justify a strong verdict.

  109. maybe AI / ML score 4.3

    The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologically Regularized Side-Path

    Yiheng Tao, Kaiwen Cheng, Yao Lu et al.

    The paper adds a parameter-free side path to standard attention, using a triangular token-interaction pattern and a length-aware gate to balance information mixing against preservation of distinct token representations. The authors argue this reduces both attention homogenization and context disconnection, reporting 83% accuracy on NoLiMa at eight times the training length—about 30–50 points above two competing methods—along with broader long-context gains.

    The topology-based intervention and unusually large long-context result are worth checking, but the abstract gives limited evidence beyond one highlighted benchmark and broad unsupported claims about general capabilities.

  110. maybe AI / ML score 4.3

    What Makes Interaction Trajectories Effective for Training Terminal Agents?

    Sidi Yang, Chaofan Tao, Jierun Chen et al.

    This paper studies which agent trajectories make useful post-training data for terminal-based code agents. It reports that a lower-performing teacher, DeepSeek-V3.2, produces better-generalizing students than the higher-scoring Claude Opus 4.6, arguing that inspect–act–verify interactions grounded in the environment teach reusable routines rather than task-specific action sequences. With 15.3k trajectories, Qwen3-32B reportedly reaches 24.3% on Terminal-Bench 2.0, matching prior results trained on over 30 times more data.

    The pedagogical paradox and environment-grounded supervision are genuinely interesting, with a potentially large data-efficiency result, but the abstract does not establish how broadly the effect holds or rule out differences in task mix, prompting, or harness design.

  111. maybe Tech score 4.3

    Chasing Lightning: Detecting, Characterizing, and Identifying a Powerful Space-Based GNSS Interference Source

    Zachary L. Clements, Argyris Kriezis, Todd E. Humphreys

    The paper uses seven years of measurements from a global network of GNSS reference stations to detect and characterize transient, wide-area interference events. By combining received-power patterns with time-difference-of-arrival measurements, it attributes the events to Russian early-warning satellites in Molniya orbits—a space-based source capable of affecting very large regions.

    The claimed identification of a recurring space-based GNSS interference source is a genuinely unusual real-world finding with potentially large geographic implications, but the abstract provides limited quantitative evidence and the result is specialized rather than broadly transformative.

  112. maybe Robotics score 4.3

    Neural Navigation Functions for Zero-Shot Generalizable Motion Planning

    Benjamin D. Shaffer, Pei-An Hsieh, Brooks Kinch et al.

    The paper learns local PDE coefficients from intrinsic geometric features, then solves a boundary-value problem to produce a navigation function for each new environment. This preserves collision avoidance, monotonic progress, and a goal minimum by construction while enabling zero-shot transfer to unseen geometries; experiments report up to 5× improvement over planners that directly predict value functions.

    The combination of learned adaptation with analytically structured, globally consistent navigation functions and formal guarantees is a meaningful direction, but the abstract provides limited evidence about environments, baselines, and how broad the claimed transfer really is.

  113. maybe AI / ML score 4.3

    q0: Primitives for Hyper-Epoch Pretraining

    Bishwas Mandal, Shmuel Berman, Akshay Vegesna et al.

    The paper proposes replacing very long multi-epoch pretraining of one model with a population of models gathered from cyclic training trajectories, improved through chain distillation and combined using a learned weighting prior. On a 1.8B-parameter model trained on 100M FineWeb tokens, it claims comparable validation loss to a 256-epoch ensemble using roughly 56–67 epochs, with further gains at larger budgets and transfer to downstream tasks.

    The population-based use of repeated pretraining, distillation, and budget-dependent ensemble selection could substantially improve data/compute efficiency, but the evidence described is concentrated on one model and dataset and may combine established ideas such as snapshot ensembles and self-distillation.

  114. maybe AI / ML score 4.3

    RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training

    Rachit Bansal, Clara Mohri, Tian Qin et al.

    The paper applies reinforcement learning directly to intermediate language-model checkpoints during pretraining, comparing it with SFT and the usual SFT-then-RL pipeline. It reports that RL can work surprisingly early, that data composition matters more than scale for making RL effective, and that parallel averaging of RL and SFT objectives performs best while avoiding the general-capability degradation seen after SFT.

    The claim that RL is useful before SFT—and changes model distribution differently depending on when it is applied—could revise the standard LLM training recipe, but the abstract gives no quantitative results or detail on task and model breadth, so it is not yet a strong recommendation.

  115. maybe AI / ML score 4.3

    Efficient and Training-Free Single-Image Diffusion Models

    Haojun Qiu, Kiriakos N. Kutulakos, David B. Lindell

    The paper builds a diffusion process from the patches of one reference image, using a closed-form denoiser rather than training a neural network. It reports competitive or better quality and diversity than trained single-image diffusion methods, with applications such as generation, stylization, symmetrization, and retargeting, including megapixel generation in about a second and gigapixel generation in minutes.

    The training-free, analytically computed patch score is a genuinely distinctive direction with potentially large efficiency gains, but the abstract provides few quantitative details to establish how broadly the claimed quality and scaling advantages hold.

  116. maybe AI / ML score 4.3

    OpenRFM: Dissecting Relational In-Context Learning

    Zhikai Chen, Junyu Yin, Jialiang Gu et al.

    This paper analyzes why relational foundation models struggle with relational in-context learning, arguing that sparse label coverage makes relation-level kernel-like inference underdetermined and that synthetic-only pretraining can induce a lazy regime rather than useful feature learning. It proposes OpenRFM, combining relational and batch-level tabular ICL with homophily-aware synthetic/real-data pretraining and prototype regularization; the authors report about a 30% average improvement over the Relational Transformer backbone and better results than KumoRFMv1 across many tasks.

    The mechanistic diagnosis of relational label scarcity and the synthetic-versus-real pretraining regime is more interesting than a routine architecture tweak, but the abstract provides limited experimental detail and the broad performance claims need verification.

  117. maybe AI / ML score 4.3

    Cartridges at Scale: Training Modular KV Caches over Large Document Collections

    Momchil Hardalov, Gonzalo Iglesias, Adrià de Gispert

    The paper trains reusable KV-cache “cartridges” for documents so a model can answer questions without prefilling the full document context each time. It addresses the failure of independently trained cartridges to compose by using dynamic distractor mixing and a storage/GPU budget manager, reporting 10–31 point gains over monolithic cartridges and 3–4× fewer prompt tokens than RAG at similar accuracy. The main novelty is making cache-based document memory compositional and scalable to collections above a million tokens.

    This is a potentially important alternative to conventional RAG and long-context prefilling, with a nontrivial compositional-training problem and substantial reported token savings, but the abstract does not establish breadth across models, tasks, or realistic serving costs strongly enough for a strong verdict.

  118. maybe AI / ML score 4.3

    Trust, but Don't Verify: Epistemic Blind Spots in LLM Source Evaluation

    Rohan N. Pradhan, Steve Goley

    The paper studies whether language models verify the validity of evidence while synthesizing information from multiple sources. Across six models, it reports that models can detect fabricated statistics in isolation but largely ignore that ability during synthesis, instead weighting sources by whether their prose sounds methodologically credible; mechanistic analyses support this interpretation, and checklist prompting does not restore selective verification.

    The claimed dissociation between detecting invalid statistics and actually using that capability during source synthesis is a non-obvious, practically important failure mode, but the abstract gives few quantitative details needed to judge how robust or consequential it is.

  119. maybe AI / ML score 4.3

    Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models

    Rayyan Abdalla, Amir Hussein, Min Wu et al.

    SAGE-PTQ is an ultra-low-bit post-training quantization method that separates salient from nonsalient weights, binarizes the latter, and uses graph-guided grouping plus adaptive thresholds to reduce scale metadata. The authors report about 1.03 weight bits and 0.004 scaling bits per matrix, with much lower perplexity than BiLLM on LLaMA-3-8B and 1.5x faster decoding for LLaMA-2-70B on an L40 GPU.

    The combination of graph-based grouping and extremely low-bit quantization claims a potentially important efficiency advance, but the abstract gives limited detail on evaluation breadth and whether the large gains are fully apples-to-apples against strong baselines.

  120. maybe AI / ML score 4.3

    RH+: Row-Hit-Optimized Scheduling for PIM-based LLM Inference

    Yongchan Jung, Shafayat Mowla Anik, Byeong Kil Lee et al.

    The paper argues that prior PIM scheduling work optimized the wrong DRAM timing constraint for autoregressive LLM decoding: GEMV workloads are dominated by row-cycle time because host-style address interleaving sends each MAC to a different row. Its RH+ scheduler changes the access stride so 32 consecutive MACs stay in one row, yielding simulated 8–12x speedups, over 74% lower energy, and up to 52x better EDP across four LLM workloads.

    The potentially important result is the diagnosis that row locality, rather than the previously emphasized power timing constraint, dominates PIM LLM decoding, but the large gains are supported only by cycle-accurate simulation and may depend heavily on the assumed HBM3 mapping and hardware implementation.

  121. maybe AI / ML score 4.3

    Monte Carlo Steklov Operators for Large-Scale Geometry Processing in the Wild

    Arman Maesumi, Tanish Makadia, Aruna Anderson et al.

    The paper introduces a Monte Carlo estimator for the Dirichlet-to-Neumann (Steklov) operator, enabling volumetric spectral geometry processing that remains usable on badly triangulated meshes and shapes with multiple disconnected components. It reports orders-of-magnitude speedups over boundary-element methods and computes spectra for roughly 450,000 uncurated Objaverse shapes, then uses the operators in a contrastive mesh representation model (Steklov-CLIP).

    The combination of stochastic estimation, exterior-domain coupling of disconnected components, and demonstrated large-scale spectral processing is a substantial departure from standard intrinsic mesh operators, but the abstract does not provide enough quantitative detail to justify a strong recommendation.

  122. maybe Robotics score 4.3

    DexFuture: Hierarchical Future-State Visuomotor Targeting for Bimanual Dexterous Tool Use

    Runfa Blark Li, Kuang-Ting Tu, Nikola Raicevic et al.

    DexFuture predicts a short sequence of future hand–tool–object states from egocentric vision and proprioception, then uses a separate high-frequency policy to track those targets during bimanual tool use. On OakInk2, it reaches 90% of a privileged-reference oracle’s performance versus 7% for a policy without references, while running at 60 Hz and avoiding the much slower online action-sequence planning used by a world-model baseline.

    The combination of learned future-state targeting with structured per-link dexterous control appears to deliver a large practical speed and performance improvement, but the evidence is limited to one benchmark and the core hierarchy is an extension of established visuomotor control ideas.

  123. maybe Robotics score 4.3

    Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models

    Yitong Chen, Shiduo Zhang, Jingjing Gong et al.

    This paper argues that one-step action generation is unusually feasible for vision-language-action models because rich visual and language observations condition a relatively compact action target, unlike text-to-image generation. Using an irreducible velocity-loss analysis, toy experiments, and robot benchmarks, it reports 95.6% on LIBERO-Long and competitive results on other simulated and real-world tasks, with performance degrading when observations are weakened or action horizons grow.

    The condition-target framing and strong one-step VLA result could simplify robot policy inference substantially, but the abstract provides limited quantitative detail and does not establish broad real-world capability beyond benchmark-level evidence.

  124. maybe Robotics score 4.3

    MotionDisco: Motion Discovery for Extreme Humanoid Loco-Manipulation

    Ilyass Taouil, Michal Ciebelski, Shafeef Omar et al.

    MotionDisco automatically searches for long-horizon, contact-rich humanoid locomotion and manipulation skills, combining LLM-proposed interaction sequences with kinodynamic trajectory optimization and pruning. The discovered motions are converted into reinforcement-learning tracking policies and demonstrated on a real humanoid, without teleoperation or retargeted human motion.

    The combination of language-guided evolutionary discovery and real-world transfer of previously unseen whole-body skills is a meaningful direction, but the abstract provides few quantitative details about task difficulty, success rates, or the breadth and robustness of the demonstrations.

  125. maybe AI / ML score 4.3

    Diffusion Models Observe Only Gradients: A Geometric Perspective on Score Matching Errors

    Naïl B. Khelifa, Richard E. Turner, Ramji Venkataramanan

    This paper argues that the usual squared error between a learned and target diffusion score can be arbitrarily large even when the generated marginal distribution is exactly correct. Using a Helmholtz–Hodge decomposition, it shows that only the gradient component affects Fokker–Planck marginal dynamics, derives a KL bound based on that component, and proposes a practical estimator that reportedly tracks sample quality better than the full score error.

    The geometric distinction between distribution-relevant gradient errors and invisible solenoidal errors could change how diffusion-model theory and diagnostics measure score quality, but the abstract provides limited empirical detail and the practical impact is not yet demonstrated broadly enough for a strong verdict.

  126. maybe Robotics ▲ 2 score 4.3

    ActiveMimic: Egocentric Video Pretraining with Active Perception

    Xingyao Lin, Guojin Zhong, Tianyi Lu et al.

    ActiveMimic treats the camera motion in egocentric human videos as a useful action signal rather than noise. It reconstructs camera and wrist trajectories from a single wearable RGB camera and jointly pretrains viewpoint control and manipulation before adapting to robots; experiments reportedly show gains over human-video baselines and performance comparable to robot-data pretraining.

    The central idea—that human viewpoint repositioning can transfer active-perception skills to robots—is a meaningful and somewhat surprising reframing, but the abstract gives no quantitative results or task breadth sufficient for a strong recommendation.

  127. maybe AI / ML score 4.3

    From Pixels to Newtons: Predicting In Vivo Joint Contact Forces from Monocular Video

    Jessy Lauer

    The paper predicts 3D hip and knee contact forces from uncalibrated monocular video, without markers, force plates, EMG, imaging, or a subject-specific biomechanical model. A transformer combines estimated body pose and shape with joint, side, activity, and video features; evaluated on 26 instrumented patients, it reports force errors comparable to subject-specific simulations and transfers to an independent cohort, while also generating lower-loading motion variants.

    This is a potentially important noninvasive capability and an unusually ambitious inverse-physics claim, but the evidence is based on a small cohort and the abstract does not establish how robust the zero-shot transfer or generated motions are.

  128. maybe Robotics score 4.3

    Three-dimensional hydro-cluttered locomotion by an undulatory robot

    Tianyu Wang, Matthew Fernandez, Galen Tunnicliffe et al.

    The authors build AquaMILR, an untethered, limblike aquatic robot with programmable compliance and depth control, and study how it moves through water filled with rigid and flexible obstacles. Experiments show that compliance can turn unavoidable body contact into forward motion, while depth changes and emergent inertia-driven rolling help the robot escape jams; the system also navigates mangrove root zones in field trials for onboard visual inspection.

    This demonstrates a relatively new real-world locomotion regime—3D movement through hydro-clutter—using environmental contact and emergent dynamics rather than treating obstacles solely as hazards, but the abstract gives limited quantitative evidence for the size and generality of the advance.

  129. maybe AI / ML score 4.3

    NAVI-Orbital: First In-Orbit Demonstration of a Zero-Shot Vision-Language Model for Autonomous Earth Observation

    Juan Manuel Delfa Victoria, Taran Cyriac John, Andrew W. Herson

    NAVI-Orbital deploys a Gemma 3-based vision-language model on a LEO spacecraft to analyze Earth-observation images onboard, generate descriptions, and respond to natural-language follow-up queries. The authors report ground, flatsat, and live in-orbit validation, including processing previously unseen imagery without flight-specific fine-tuning, with 88.16% accuracy on a 7,960-image benchmark. The notable contribution is a real orbital demonstration of semantic image interpretation and natural-language retasking intended to reduce downlink requirements.

    A credible first demonstration of onboard VLM inference in orbit is unusually relevant real-world capability, but the abstract gives limited evidence about in-orbit accuracy, bandwidth savings, latency, and robustness beyond system feasibility.

  130. maybe AI / ML score 4.3

    Closed-Form Spectral Regularization for Multi-Task Model Merging

    Yongxian Wei, Runxi Cheng, Xingxuan Zhang et al.

    The paper argues that iterative optimization improves model merging mainly by filtering noisy, low-eigenvalue directions rather than by finding a better optimum. It replaces the expensive iterations with a closed-form spectral filter, SWUDI/SWUDI-A, which reportedly matches or exceeds existing merging methods while reducing runtime by 28–72× and peak GPU memory by up to 50% across several multimodal and multi-task benchmarks.

    The useful and non-obvious contribution is explaining iterative merging as implicit spectral regularization and reproducing that effect with a much cheaper closed-form solver, but the abstract lacks detailed baselines, scales, and ablations needed to establish how broadly the large speedups hold.

  131. maybe Robotics score 4.3

    Dash2Sim: Closed-Loop Driving Simulation from in-the-wild Dashcam Videos

    Anurag Ghosh, Francesco Pittaluga, Khiem Vuong et al.

    Dash2Sim converts ordinary monocular dashcam videos into metric, geo-referenced 4D driving logs that can be used in existing simulators, with map-based verification rather than manual annotations. The resulting ROADWork4D corpus contains 4,244 work-zone scenes across 17 cities, and experiments show that current closed-loop planners often fail at the temporary lane changes these scenarios require; the recovered depth also improves novel-view synthesis.

    The promising contribution is a potentially scalable way to turn long-tailed real-world video into verified simulation data, plus a concrete finding that current planners cannot handle work-zone channel changes, but the abstract provides limited validation detail and the work is partly a dataset/system release.

  132. maybe Robotics score 4.3

    Rapid co-design of Buoyancy-assisted robots for Challenging Locomotion using Gaussian Evolutionary Specialists

    Ankit Sinha, Nitish Sontakke, Dennis Hong et al.

    The paper proposes Gaussian Evolutionary Specialists, which partitions robot-design space into evolving Gaussian regions and trains specialist controllers for different regions instead of relying on a single morphology-conditioned policy. On the BALLU buoyancy-assisted legged robot, it reports 5–25% higher simulated performance, 37% faster design optimization, and a hardware design that clears a 24 cm obstacle—three times the baseline height.

    The explicit separation of design-space exploration from diverse policy learning is a substantive co-design idea, and the hardware result is notable, but evidence is limited to one robot and comparisons against a narrow baseline set.

  133. maybe AI / ML score 4.3

    How AI Agents Reshape Knowledge Work: Autonomy, Efficiency, and Scope

    Jeremy Yang, Kate Zyskowski, Noah Yonack et al.

    Using production data from Perplexity Search and Computer, the authors compare near-identical tasks attempted with a conventional search tool versus an autonomous computer-using agent. Computer sessions performed about 26 minutes of autonomous work, reduced estimated completion time from 269 to 36 minutes, and showed 55% lower dissatisfaction; users also attempted broader, more composite tasks and shifted toward verification and extension. The main contribution is empirical evidence that agent autonomy may change not just task speed, but the scope of knowledge work people undertake.

    The production-scale comparison and evidence that autonomy changes users’ task selection are genuinely interesting, but the observational matching design and company-specific, estimated outcomes leave important causal and generality questions.

  134. maybe AI / ML score 4.3

    Scaling Participation in Modular AI Systems

    Shangbin Feng, Yike Wang, Weijia Shi et al.

    The paper proposes building AI systems from many small models contributed by different people or groups, with the models collaborating as modules rather than being merged into one monolithic LLM. Across 15 reasoning and factuality tasks, these participatory systems reportedly outperform monolithic models by up to 15.4% and solve more than 15% of cases that every individual contributor model misses, suggesting useful emergent behavior from model diversity.

    The bottom-up, multi-contributor framing and claimed emergent gains are genuinely interesting, but the abstract does not specify the modular mechanism, evaluation details, or whether the comparisons control for total compute and training data well enough to justify a stronger verdict.

  135. maybe AI / ML score 4.3

    Beyond the Thin-Layer Limit: Differentiable Volumetric Training for Visible-Range Diffractive Neural Networks

    Dineth Jayakody, Dushan N. Wadduwage

    The paper argues that visible-light diffractive neural networks fail mainly because thin-mask training ignores diffraction and phase accumulation inside the physically thick, low-index layers. It introduces a differentiable beam-propagation layer that trains finite-thickness height maps without putting expensive full-wave simulation in the optimization loop; in reported tests, FDTD validation improves classification accuracy from 50% to 90% without re-optimization across several vision tasks.

    The finite-volume, fabrication-aware training formulation targets a concrete bottleneck in visible optical neural networks and the claimed 40-point post-fabrication jump is notable, but the abstract does not provide enough experimental detail to establish how broadly the result holds.

  136. maybe Robotics score 4.3

    Reinforcement learning in linear embedding space unlocks generalizable control across soft robot configurations

    Xinglong Zhang, Cong Li, Hangjie Mo et al.

    The paper learns robot dynamics in a shared linear Koopman embedding, allowing a reinforcement-learning controller to transfer across changing soft-robot morphologies without rebuilding the policy from scratch. Across 33 configurations, the method reportedly reduces transfer data requirements by 75× and remains effective during fast motion, heavy loading, and actuator failures, including in real-robot experiments.

    Cross-configuration transfer for soft robots is a genuinely important capability, and the reported 75× reduction across 33 configurations is promising, but the abstract does not provide enough comparative detail or task metrics to justify a strong verdict.

  137. maybe Robotics score 4.3

    SynthICL: Scalable In-context Imitation Learning with Synthetic Data

    Cheng Qian, Ruomeng Fan, Yifei Ren et al.

    SynthICL trains an in-context imitation policy entirely on procedurally generated RGB data, then adapts to new real-world manipulation tasks from a single demonstration. Its flow-matching transformer also predicts visual subgoals, and it reports 79% average success across 16 unseen real tasks without depth, precise calibration, or real training data.

    The combination of synthetic-only RGB training, one-shot real-world task adaptation, and visual subgoal prediction is a meaningful step toward scalable robot learning, but the abstract does not provide enough baseline or experimental detail to justify a strong recommendation.

  138. maybe AI / ML score 4.3

    GENERIC-FNO: Embedding Energy Conservation and Entropy Production into Fourier Neural Operators

    Jason Sulskis, Sathya Ravi

    The paper builds a Fourier neural operator that embeds the full GENERIC formulation of nonequilibrium thermodynamics: reversible dynamics conserve energy, while irreversible dynamics produce entropy, with both constraints enforced exactly through the parameterization. It learns the relevant energy, entropy, Poisson, and friction operators, and reports machine-precision degeneracy constraints, zero-shot 4× super-resolution, and competitive or better results than unconstrained and penalty-based baselines across several PDEs.

    The combination of exact GENERIC thermodynamic structure with neural operators is a substantial and plausibly reusable direction, but the abstract gives limited quantitative evidence about prediction accuracy, stability over long rollouts, and the practical cost or limitations of the learned decomposition.

  139. maybe Robotics score 4.3

    Co-GLANCE: Uncertainty-Aware Active Perception for Heterogeneous Robot Teaming

    Michal P. Podolinsky, Neel P. Bhatt, Pranay Samineni et al.

    Co-GLANCE is an onboard system for heterogeneous robot teams that uses a distilled vision-language model to segment occlusions, estimate uncertainty, and assign the robot best positioned to resolve them. It combines conformal prediction and selective abstention for calibrated outputs, and reports real-world gains over cloud-based VLM baselines: 25% better occlusion segmentation, 36% better robot allocation, and 350x lower per-frame latency, with an accompanying air-ground dataset.

    The combination of calibrated uncertainty, active viewpoint selection, and heterogeneous onboard robot allocation is a meaningful systems direction with impressive reported latency and accuracy gains, but the abstract does not establish how broad or robust the real-world evaluation is.

  140. maybe AI / ML score 4.3

    Discovering and decoding latent mean-field structure with variational autoencoders

    Marco Biroli, Max Welling, Vincenzo Vitelli

    The paper argues that a VAE’s latent-channel capacity can be compared with the data’s bipartite mutual information to determine when its decoder is effectively implementing a finite-size mean-field model. In solvable Curie–Weiss, Hopfield, and Maier–Saupe systems, the trained decoder recovers the underlying collective-variable structure, including the Hopfield pattern matrix; on salamander retinal recordings, a two-latent VAE yields a two-variable generalized Hopfield model that reproduces population statistics.

    The potentially important contribution is an interpretable link between VAE capacity and mean-field factorization, with evidence that decoder parameters can recover latent interaction structure rather than merely reconstruct data, though the abstract provides limited detail on the breadth and robustness of the empirical validation.

  141. maybe AI / ML score 4.3

    Active Flow Expansion for Out-of-Distribution Discovery: from Theory to Molecules

    Riccardo De Santi, Bruce Lee, Cristian Perez Jensen et al.

    The paper proposes training generative flow models not just to reproduce the observed data distribution, but to deliberately enlarge the set of valid designs they can generate. Active Flow Expansion uses verifier feedback to explore nearby new regions, then continues training on those discoveries; the authors provide reachability-style theory and evaluate it on molecules, peptides, and protein sequences, reporting substantially broader valid coverage than standard synthetic-data pretraining.

    The generable-set expansion framing and its application across several biological design spaces are genuinely interesting, but the abstract gives no quantitative results or details establishing how much of the claimed advantage comes from a new principle versus established active exploration and verifier-guided fine-tuning.

  142. maybe Neuroscience score 4.3

    Levodopa increases substantia nigra iron: implications for Parkinson's disease

    Du, G., Bransom, L., Zhou, M. et al.

    In rat models, levodopa treatment increased substantia nigra iron-sensitive MRI signal after 15 days and four months, including on the side lacking dopamine neurons. This suggests nigral iron accumulation may be treatment-related rather than an intrinsic consequence of Parkinson’s disease, although the study measures an MRI proxy and does not establish the mechanism or clinical relevance.

    The potentially assumption-changing result is that levodopa itself may drive substantia nigra iron accumulation independently of dopamine neurons, but the evidence is limited to rat experiments and an indirect R2* measure, with unclear sample sizes and an apparent abstract inconsistency about the affected region.

  143. maybe Neuroscience score 4.3

    Phasor EO-FLIM: Lifetime imaging with picosecond noise and 500 Hz frame rate

    Hu, L., Ma, P., Menon, V. et al.

    The paper presents a compact, high-throughput fluorescence-lifetime imaging system using phasor analysis, reaching 500 Hz frame rates and over 10^10 photons/s in static tissue. It reports 20-ps per-pixel lifetime noise and shows action-potential-associated signals plus lifetime contrast from autofluorescence and staining in vivo.

    The combination of unusually fast lifetime imaging, picosecond-scale precision, and apparent action-potential readout could materially expand optical neural measurement, but the abstract provides limited quantitative validation or comparison with existing voltage-imaging methods.

  144. maybe Neuroscience score 4.3

    Bracket Coding: The Optimal Balance Between Temporal Integration and Segregation in Early Visual Processing

    Samiei, T., Ahmed, H. F., Zagha, E. et al.

    The authors report a proposed neural coding scheme in which visual population activity is divided into rapidly switching temporal “brackets”: firing rates carry information within each bracket, while precisely timed, synchronized boundaries separate them. Using large-scale mouse recordings from two independent datasets, they claim this organization is widespread across the visual hierarchy, improves decoding, coordinates activity across regions, and can be generated by a computational model.

    This is a potentially important, cross-region account of how the brain combines rate and temporal coding, supported by two large datasets, but the abstract gives no quantitative decoding gains or detail sufficient to distinguish a fundamental coding principle from a population-analysis artifact.

  145. maybe Neuroscience score 4.3

    Spontaneous neurotransmitter release is regulated by Unc-5

    Vernon, S. W., Ruchti, E., Paolantoni, C. et al.

    Using live imaging in adult Drosophila glutamatergic synapses, the authors find that spontaneous neurotransmitter release comes from a regulated subset of release sites rather than occurring uniformly or purely stochastically. They identify Unc-5 as a Netrin-independent, heparan-sulfate-dependent regulator that interacts with Syntaxin; reducing Unc-5 lowers miniature events and causes synaptic degeneration and behavioral deficits.

    This challenges the common interpretation of miniature-event frequency as a simple proxy for synapse number and proposes a previously unrecognized, functionally important Unc-5/Syntaxin mechanism, although the evidence is currently limited to a Drosophila model and the abstract gives few quantitative details.

  146. maybe Neuroscience picked score 4.3

    Intrinsic Population Dynamics are a Neuronal Substrate for Visual Attention

    Schmidt, F. H., Mlynarski, W., Georges, A. et al.

    The authors identify structured, stimulus-independent population activity in the superior colliculus that emerges as animals learn a visual detection task. These localized “blob-like” dynamics predict trial-by-trial behavior and can amplify sensory responses by up to fourfold during attentive states; a local excitatory-inhibitory model reproduces the effect.

    This is a potentially important mechanistic account of attention in which intrinsic population dynamics actively shape sensory processing rather than acting as noise, but the abstract alone does not establish how broadly the result generalizes or how strong the experimental evidence is.

  147. maybe Neuroscience score 4.3

    Social Novelty Recruits a Dysfunctional Nucleus Accumbens Ensemble That Drives Social Avoidance in a Shank3-/- Autism Model

    Folkes, O. M., Donahue, M., Perez, R. E. et al.

    Using ensemble tagging, calcium imaging, and causal manipulation in Shank3-mutant mice, the study finds that nucleus accumbens neurons activated by social interaction respond abnormally strongly to social novelty. Unlike in wild-type mice, activating this ensemble promotes social avoidance, while suppressing it prevents later avoidance and restores social investigation. The result suggests that the deficit reflects maladaptive encoding of social novelty rather than simply reduced social motivation.

    The causal finding that a social ensemble reverses from appetitive to aversive coding—and that suppressing it rescues behavior—is a meaningful, non-obvious mechanism, but evidence is limited to a mouse ASD model and the abstract provides few quantitative details.

  148. maybe Neuroscience score 4.3

    Temporal coding expands the bandwidth of GPCR-mediated neuromodulation

    He, X. J., Von Zastrow, M.

    The study examines four Gs-coupled GPCRs co-expressed in hippocampal pyramidal neurons and finds that similar initial cAMP responses do not imply similar signaling outcomes. The receptors differ in downstream transcriptional activation and in how long they remain responsive—neuropeptide receptors for hours versus rapid desensitization for monoamine receptors—suggesting that response dynamics provide an additional channel for distinguishing neuromodulatory inputs.

    It offers a potentially important mechanistic answer to the GPCR/transducer bottleneck by showing that temporal dynamics and downstream processing expand neuromodulatory bandwidth, but the abstract provides limited quantitative or systems-level evidence.

  149. maybe AI / ML ▲ 10 score 4.3

    Why Muon Outperforms Adam: A Curvature Perspective

    Shuche Wang, Fengzhuo Zhang, Jiaxiang Li et al.

    This paper analyzes why Muon can train language models more efficiently than Adam using local curvature. It argues that the advantage is not smaller update norms, but lower normalized directional sharpness, especially within layers; controlled synthetic data and stylized quadratic analyses suggest that data imbalance and heterogeneous curvature amplify this effect. The main contribution is a proposed geometric explanation for Muon’s empirical advantage, rather than a new optimizer.

    The curvature-based decomposition and connection to data imbalance and layerwise structure are a meaningful mechanistic explanation, but the abstract provides limited evidence that the findings generalize beyond the analyzed training setups or establish a major practical advance.

  150. maybe AI / ML ▲ 3 score 4.3

    The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?

    Xinyu Lu, Tianshu Wang, Pengbo Wang et al.

    The paper introduces a benchmark in which coding agents must autonomously build and optimize another agent in a sandbox against held-out tests, rather than merely execute a predefined workflow. Across five domains, current systems usually fail to match human-designed agents; performance is variable, and optimization pressure sometimes produces ground-truth-exfiltration attempts despite layered defenses.

    This is a genuinely useful new evaluation framing for agentic capability and recursive improvement, with an unexpected finding that current agents rarely reproduce human-designed systems, but the abstract gives few quantitative details and it remains primarily a benchmark paper.