<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
  <title>Paper Feed</title>
  <link>https://paperfeed.app/</link>
  <atom:link href="https://paperfeed.app/feed.xml" rel="self" type="application/rss+xml"/>
  <description>Weekly picks of genuinely interesting research papers, with explainers.</description>
  <lastBuildDate>Sun, 30 Aug 2026 20:55:25 +0000</lastBuildDate>
  <item>
    <title>Issue 34 · Aug 17–23, 2026</title>
    <link>https://paperfeed.app/2026-W34/index.html</link>
    <guid isPermaLink="true">https://paperfeed.app/2026-W34/index.html</guid>
<pubDate>Sun, 30 Aug 2026 05:33:30 +0000</pubDate>    <description>This week clusters around two big themes. First, a wave of results questioning whether our models represent what we think they do: causal experiments show that near-identical neural predictivity does not mean brain-like representations, and a related paper finds brain alignment appears before any learning. Second, a striking security thread on LLMs quietly leaking secrets—from context, from hidden reasoning traces, and even from memorized data unlocked by innocuous fine-tuning. Alongside these are genuinely new capability demonstrations in robotics (policies that self-improve from their own failures) and neuroscience (monkeys trained to report their own cortical activity). Treat the strongest benchmark claims as upper bounds until you read the methods. • 1. Parametric neural control differentiates top neural network models of primate visual cortex • 2. Reinforcement Learning on Benign Facts Amplifies Leakage of Memorized Private Data • 3. Monkeys learn to report their own sensory cortical population activity • 4. Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning • 5. Inductively Scalable, Single-Step Neural Surrogates for Wave-Scattering Inverse Problems • 6. Efficient coding makes and breaks Webers law</description>
  </item>
  <item>
    <title>1. Parametric neural control differentiates top neural network models of primate visual cortex</title>
    <link>https://paperfeed.app/2026-W34/biorxiv-10-64898-2026-08-16-745063.html</link>
    <guid isPermaLink="false">paperfeed:2026-W34:biorxiv:10.64898/2026.08.16.745063</guid>
<pubDate>Sun, 30 Aug 2026 05:33:30 +0000</pubDate>    <description>A causal, closed-loop test showing that vision models with indistinguishable neural predictivity diverge sharply in their ability to actually drive the neurons they claim to model—directly puncturing the assumption that predictivity implies a shared brain-aligned parameterization. The scale (27,500 stimuli, five macaques, ten models, multiple visual areas) and the finding that input-gradient spatial-frequency structure predicts control better than accuracy make this a rare assumption-breaking result at the AI–neuroscience interface. Look for: Check how &#39;control&#39; is operationalized versus in-distribution predictivity, and whether the adversarial-training advantage is confounded with the gradient-spectrum predictor they favor. (Prince, J. S., Wang, B., Fel, T. et al., https://www.biorxiv.org/content/10.64898/2026.08.16.745063v1)</description>
  </item>
  <item>
    <title>2. Reinforcement Learning on Benign Facts Amplifies Leakage of Memorized Private Data</title>
    <link>https://paperfeed.app/2026-W34/arxiv-2608-21727.html</link>
    <guid isPermaLink="false">paperfeed:2026-W34:arxiv:2608.21727</guid>
<pubDate>Sun, 30 Aug 2026 05:33:30 +0000</pubDate>    <description>A genuinely counterintuitive privacy failure: reinforcement learning on benign factual data that contains no private information makes a model surface PII it had memorized earlier, with a 2.4x jump in verbatim recall on DeepSeek-V3.1 and the effect growing with scale. It reframes memorized-data extraction as something an adversary can unlock without ever touching the data, which matters for anyone fine-tuning released models. Look for: Whether the &#39;memorized but latent&#39; baseline is measured cleanly and whether the effect is specific to RLVR or would appear under ordinary SFT too. (Renfei Zhang, Niloofar Mireshghallah, https://arxiv.org/abs/2608.21727)</description>
  </item>
  <item>
    <title>3. Monkeys learn to report their own sensory cortical population activity</title>
    <link>https://paperfeed.app/2026-W34/biorxiv-10-64898-2026-08-07-743408.html</link>
    <guid isPermaLink="false">paperfeed:2026-W34:biorxiv:10.64898/2026.08.07.743408</guid>
<pubDate>Sun, 30 Aug 2026 05:33:30 +0000</pubDate>    <description>Using online V4 recordings and closed-loop feedback, macaques learned to base decisions on specific axes of their own population activity rather than on the stimulus—a direct causal probe of sensory readout that is close in spirit to brain-computer interfaces. The result that stimulus–choice misalignment reflects the available training signal rather than an intrinsic readout limit is a meaningful shift in how to think about perceptual decision-making. Look for: Scrutinize the controls ruling out changes in stimulus selectivity/noise correlations, and how much of the learned readout is genuinely novel versus reweighting of existing variance. (Hu, J., Okazawa, G., https://www.biorxiv.org/content/10.64898/2026.08.07.743408v1)</description>
  </item>
  <item>
    <title>4. Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning</title>
    <link>https://paperfeed.app/2026-W34/arxiv-2608-21204.html</link>
    <guid isPermaLink="false">paperfeed:2026-W34:arxiv:2608.21204</guid>
<pubDate>Sun, 30 Aug 2026 05:33:30 +0000</pubDate>    <description>A simple, scalable recipe for making large imitation-learned robot policies learn from their own deployment failures: attach a small off-policy Q-function, keep the billion-parameter BC policy frozen, and reweight/finetune only the critic. Real bimanual gains (cup stacking 40%→90%, wallet insertion 25%→80%) on contact-rich tasks without new human demos address a core limitation of behavior cloning. Look for: How the Q-function avoids overestimation on self-generated failures, and whether gains persist beyond the near-ceiling simulated suites into harder real tasks. (Varun Giridhar, Anant Khandelwal, Jeremy A. Collins et al., https://arxiv.org/abs/2608.21204)</description>
  </item>
  <item>
    <title>5. Inductively Scalable, Single-Step Neural Surrogates for Wave-Scattering Inverse Problems</title>
    <link>https://paperfeed.app/2026-W34/arxiv-2608-17344.html</link>
    <guid isPermaLink="false">paperfeed:2026-W34:arxiv:2608.17344</guid>
<pubDate>Sun, 30 Aug 2026 05:33:30 +0000</pubDate>    <description>Single-step neural surrogates for electromagnetic wave scattering have been stuck at tens of variables; by actively generating training examples where the surrogate disagrees most with a full-wave solver, this scales to ~42k trainable variables and generalizes inductively to over 3 million—a real jump on a bottleneck in neural physical simulation, with concrete photonic inverse-design demonstrations. Look for: Whether inductive generalization to 3M+ variables holds accuracy across diverse structures, and how the FDTD speedups are measured (range varies widely, 1.29–26.5x). (Charles Dove, Laura Waller, https://arxiv.org/abs/2608.17344)</description>
  </item>
  <item>
    <title>6. Efficient coding makes and breaks Webers law</title>
    <link>https://paperfeed.app/2026-W34/biorxiv-10-64898-2026-08-10-744043.html</link>
    <guid isPermaLink="false">paperfeed:2026-W34:biorxiv:10.64898/2026.08.10.744043</guid>
<pubDate>Sun, 30 Aug 2026 05:33:30 +0000</pubDate>    <description>Direct causal evidence that Weber&#39;s law is not a fixed property of perception but emerges from efficient coding: skewing the stimulus distribution toward large magnitudes inverts the usual discriminability pattern across three sensory modalities. Replacing a descriptive psychophysical regularity with a mechanistic, distribution-dependent explanation is exactly the kind of result that reshapes how we think about perception. Look for: Effect sizes and whether the inversion is robust across individuals and modalities, and how well efficient-coding predictions match the quantitative adaptation. (Prat-Carrabin, A., Yamamoto, R., Gershman, S. J., https://www.biorxiv.org/content/10.64898/2026.08.10.744043v1)</description>
  </item>
  <item>
    <title>Project 1: moonshotai/Kimi-K3 (HF model)</title>
    <link>https://paperfeed.app/2026-W34/hf-model-moonshotai-Kimi-K3.html</link>
    <guid isPermaLink="false">paperfeed:2026-W34:hf:model/moonshotai/Kimi-K3</guid>
<pubDate>Sun, 30 Aug 2026 05:33:30 +0000</pubDate>    <description>A 2.8T-parameter sparse MoE with native vision, 1M-token context, and several real architectural innovations (Kimi Delta Attention, Attention Residuals, Stable LatentMoE) is a rare open frontier release rather than a fine-tune. If the evals hold, this is the most consequential open model of the week and a template others will study. Look for: Verify the reported benchmarks and that the released weights actually match the described 104B-active/896-expert architecture; serving it is nontrivial. (https://huggingface.co/moonshotai/Kimi-K3)</description>
  </item>
  <item>
    <title>Project 2: Anthropic/claude-protein-binder-design (HF dataset)</title>
    <link>https://paperfeed.app/2026-W34/hf-dataset-Anthropic-claude-protein-binder-design.html</link>
    <guid isPermaLink="false">paperfeed:2026-W34:hf:dataset/Anthropic/claude-protein-binder-design</guid>
<pubDate>Sun, 30 Aug 2026 05:33:30 +0000</pubDate>    <description>An openly released set of 1,440 de novo miniprotein binders designed by autonomous Claude agents, with 354 experimentally validated binders, raw assay data, structures, and full provenance is a genuinely new research artifact at the AI-for-science frontier. It lets you study where LLM-driven design actually succeeds or fails rather than trusting a headline. Look for: Check success rates per target and whether binding was independently confirmed; this is a data release, not a reusable design model. (https://huggingface.co/datasets/Anthropic/claude-protein-binder-design)</description>
  </item>
  <item>
    <title>Project 3: MiniMaxAI/MiniMax-H3 (HF model)</title>
    <link>https://paperfeed.app/2026-W34/hf-model-MiniMaxAI-MiniMax-H3.html</link>
    <guid isPermaLink="false">paperfeed:2026-W34:hf:model/MiniMaxAI/MiniMax-H3</guid>
<pubDate>Sun, 30 Aug 2026 05:33:30 +0000</pubDate>    <description>Unified reference-conditioned generation of up to 15s of 2K video with synchronized 32kHz stereo audio from text/image/video/audio inputs is a real capability jump for omni-modal generation. The heavy ecosystem of derived Spaces, LoRAs, and workflows this week signals it is already shaping how people build AV pipelines. Look for: The crucial Context-IR preprocessing remains hosted/closed, so the release is only partially open and true reproducibility is limited. (https://huggingface.co/MiniMaxAI/MiniMax-H3)</description>
  </item>
  <item>
    <title>Project 4: Qwen/Qwen3.8-Flash-Next (HF model)</title>
    <link>https://paperfeed.app/2026-W34/hf-model-Qwen-Qwen3-8-Flash-Next.html</link>
    <guid isPermaLink="false">paperfeed:2026-W34:hf:model/Qwen/Qwen3.8-Flash-Next</guid>
<pubDate>Sun, 30 Aug 2026 05:33:30 +0000</pubDate>    <description>An experimental preview of Qwen4&#39;s architecture combining Gated DeltaNet recurrence, block-level sparse attention, gated residuals, and large offload-friendly n-gram embeddings represents several substantive departures from standard transformer scaling. At 125B/6B-active it is runnable and gives an early look at where long-context efficiency is heading. Look for: Practical long-context gains are vendor-reported; validate independently, and note it is an experimental preview rather than a stable release. (https://huggingface.co/Qwen/Qwen3.8-Flash-Next)</description>
  </item>
  <item>
    <title>Project 5: FireRedTeam/FireRedAudio (GitHub)</title>
    <link>https://paperfeed.app/2026-W34/github-FireRedTeam-FireRedAudio.html</link>
    <guid isPermaLink="false">paperfeed:2026-W34:github:FireRedTeam/FireRedAudio</guid>
<pubDate>Sun, 30 Aug 2026 05:33:30 +0000</pubDate>    <description>A 9B audio-language model that unifies ASR, audio reasoning, hour-long temporal grounding, zero-shot/instruct TTS, voice design, and semantic/acoustic speech editing via decoupled understanding/generation pathways is meaningfully broader than a typical TTS or ASR system. The decoupled continuous-representation design with a shared backbone is an interesting architectural bet with released code and weights. Look for: Benchmark-leadership claims are thinly evidenced in the card and the HF weights show near-zero downloads; test the editing and grounding quality yourself. (https://github.com/FireRedTeam/FireRedAudio)</description>
  </item>
  <item>
    <title>Project 6: ShareLab-SII/VA-Judger (GitHub)</title>
    <link>https://paperfeed.app/2026-W34/github-ShareLab-SII-VA-Judger.html</link>
    <guid isPermaLink="false">paperfeed:2026-W34:github:ShareLab-SII/VA-Judger</guid>
<pubDate>Sun, 30 Aug 2026 05:33:30 +0000</pubDate>    <description>Billed as the first reward model for joint video-audio generation, it models cross-modal semantic and temporal coherence via pairwise human preferences plus dimension-wise rewards, with released checkpoints, a benchmark, and a runnable RL post-training LoRA. It targets a central unsolved problem in AV generation rather than tweaking a metric. Look for: Verify the reward model generalizes beyond the LTX-2 setup and that the preference benchmark correlates with human judgments at scale. (https://github.com/ShareLab-SII/VA-Judger)</description>
  </item>
  <item>
    <title>Issue 29 · Jul 13–19, 2026</title>
    <link>https://paperfeed.app/2026-W29/index.html</link>
    <guid isPermaLink="true">https://paperfeed.app/2026-W29/index.html</guid>
<pubDate>Sun, 30 Aug 2026 20:51:58 +0000</pubDate>    <description>This week is heavy on assumption-breaking audits and genuinely new training directions. Two threads stand out: papers that show widely-used recipes or metrics measure something other than we assume (distributional-RL risk heads, answer-conditioned distillation, FDR control), and papers opening new axes of capability (test-time training as robot memory, permissionless internet-scale pretraining, a neuromorphic in-vivo closed loop). Robotics, neuroscience/BCI, and real-time video generation are all represented; treat the flashier technical-report claims with the usual skepticism about baselines and generality. • 1. RoboTTT: Context Scaling for Robot Policies • 2. Low-latency neuromorphic closed-loop control of hippocampal ripples in vivo • 3. The Benjamini--Hochberg Procedure Can Fail to Control the FDR for Correlated Two-Sided Gaussian Tests • 4. Agora: Collective and Permissionless Internet-Scale Pretraining of Large Language Models • 5. Auditing the Risk Claims of Distributional Reinforcement Learning • 6. Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models • 7. FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers • 8. Toward a mechanistic understanding of inference in visual cortex and diffusion models • 9. Towards Human-level Dexterous Teleoperation • 10. Verbalizable Representations Form a Global Workspace in Language Models</description>
  </item>
  <item>
    <title>1. RoboTTT: Context Scaling for Robot Policies</title>
    <link>https://paperfeed.app/2026-W29/arxiv-2607-15275.html</link>
    <guid isPermaLink="false">paperfeed:2026-W29:arxiv:2607.15275</guid>
<pubDate>Sun, 30 Aug 2026 20:51:58 +0000</pubDate>    <description>Uses test-time training (gradient-updated fast weights as recurrent memory) to scale robot visuomotor context to 8K steps without added latency, unlocking one-shot imitation from human video, online improvement, and a ten-stage assembly task no baseline completes. This is a genuinely new context-scaling mechanism for robot foundation models, not a variant, with a reported first-of-its-kind pretraining-context scaling law. Look for: Check how broadly the real-robot results generalize beyond the demonstrated tasks and whether the 8K-context gains hold outside the specific manipulation setups. (Yunfan Jiang, Yevgen Chebotar, Ruijie Zheng et al., https://arxiv.org/abs/2607.15275)</description>
  </item>
  <item>
    <title>2. Low-latency neuromorphic closed-loop control of hippocampal ripples in vivo</title>
    <link>https://paperfeed.app/2026-W29/biorxiv-10-64898-2026-07-09-737518.html</link>
    <guid isPermaLink="false">paperfeed:2026-W29:biorxiv:10.64898/2026.07.09.737518</guid>
<pubDate>Sun, 30 Aug 2026 20:51:58 +0000</pubDate>    <description>A 41-neuron spiking network on SpiNNaker detects hippocampal sharp-wave ripples and triggers optogenetic inhibition in awake mice at ~50 ms latency using up to 200x less energy than deep models. This is a rare fully integrated neuromorphic sense-to-stimulate loop fast enough to manipulate transient neural events causally, directly relevant to BCI and closed-loop neuromodulation. Look for: Scrutinize detection accuracy versus deep baselines and whether the causal effect on ripple dynamics is robust across the 23 sessions rather than a subset. (Alves, P., Jurado-Parras, M.-T., Freitas, J. et al., https://www.biorxiv.org/content/10.64898/2026.07.09.737518v1)</description>
  </item>
  <item>
    <title>3. The Benjamini--Hochberg Procedure Can Fail to Control the FDR for Correlated Two-Sided Gaussian Tests</title>
    <link>https://paperfeed.app/2026-W29/arxiv-2607-12208.html</link>
    <guid isPermaLink="false">paperfeed:2026-W29:arxiv:2607.12208</guid>
<pubDate>Sun, 30 Aug 2026 20:51:58 +0000</pubDate>    <description>Constructs a correlated Gaussian factor model with a rigorous interval-arithmetic certificate showing Benjamini-Hochberg exceeds its nominal FDR, disproving a 20-year-old conjecture about two-sided Gaussian tests under dependence. Foundational, surprising, and notable additionally because the proof was generated by GPT-5.6 Pro and human-checked. Look for: The effect size is tiny (FDR&gt;0.0104 at alpha=0.01); read for whether this is a knife-edge counterexample or a practically meaningful failure regime. (Edgar Dobriban, https://arxiv.org/abs/2607.12208)</description>
  </item>
  <item>
    <title>4. Agora: Collective and Permissionless Internet-Scale Pretraining of Large Language Models</title>
    <link>https://arxiv.org/abs/2607.13332</link>
    <guid isPermaLink="false">paperfeed:2026-W29:arxiv:2607.13332</guid>
<pubDate>Sun, 30 Aug 2026 20:51:58 +0000</pubDate>    <description>Demonstrates permissionless, collectively-owned pretraining of an 8.6B model on 500B tokens across 330 heterogeneous internet-connected GPUs, reportedly reaching 63% of a centralized H100 baseline with similar convergence. If the efficiency and convergence claims hold, this is a genuinely new direction for who can train frontier-scale models. Look for: Independent validation is thin; watch for how the 63% efficiency was measured and whether convergence truly matches centralized runs at this scale. (Gil Avraham, Violetta Shevchenko, Hadi Mohaghegh Dolatabadi et al., https://arxiv.org/abs/2607.13332)</description>
  </item>
  <item>
    <title>5. Auditing the Risk Claims of Distributional Reinforcement Learning</title>
    <link>https://arxiv.org/abs/2607.11607</link>
    <guid isPermaLink="false">paperfeed:2026-W29:arxiv:2607.11607</guid>
<pubDate>Sun, 30 Aug 2026 20:51:58 +0000</pubDate>    <description>Audits whether distributional-RL agents&#39; learned return distributions actually support the risk-sensitive claims read from them, and finds 40-95% of the strongest risk trade-offs are statistically false and structural artifacts, replicated to Atari scale with strong controls. This directly challenges a widespread interpretation used for interpretability and safety monitoring. Look for: The positive controls are what make this convincing—check that the audit&#39;s harness genuinely detects real effects and isn&#39;t over-rejecting due to its own noise model. (Hari Prasad, https://arxiv.org/abs/2607.11607)</description>
  </item>
  <item>
    <title>6. Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models</title>
    <link>https://arxiv.org/abs/2607.14552</link>
    <guid isPermaLink="false">paperfeed:2026-W29:arxiv:2607.14552</guid>
<pubDate>Sun, 30 Aug 2026 20:51:58 +0000</pubDate>    <description>Shows the common distillation fix of showing a model the gold answer and asking it to write a CoT degrades verifiable-reasoning accuracy by up to ~27 points, because the traces rationalize backward, and this harm is detectable from unlabeled traces and transfers across teacher families. A non-obvious failure in a widely-used recipe with an actionable &#39;generate answer-blind&#39; takeaway. Look for: Verify the controlled ablation isolating the &#39;rationalize-toward&#39; instruction from mere answer visibility, and whether the effect persists with modern filtering beyond correctness checks. (Jungseob Lee, Seungyoon Lee, Suhyune Son et al., https://arxiv.org/abs/2607.14552)</description>
  </item>
  <item>
    <title>7. FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers</title>
    <link>https://arxiv.org/abs/2607.14898</link>
    <guid isPermaLink="false">paperfeed:2026-W29:arxiv:2607.14898</guid>
<pubDate>Sun, 30 Aug 2026 20:51:58 +0000</pubDate>    <description>Replaces the slow 3D-convolutional video latent decoder with a causal Transformer using a bounded rolling KV cache, decoding pixels frame-by-frame at 3.6-4.7x speed and up to 11x less memory at 1080p while matching reconstruction quality. Targets a real and under-addressed bottleneck for real-time video generation. Look for: Evidence is limited to two Wan latent spaces on one GPU; look for whether quality holds for long videos and complex motion, not just PSNR at 1080p. (Minguk Kang, Suha Kwak, https://arxiv.org/abs/2607.14898)</description>
  </item>
  <item>
    <title>8. Toward a mechanistic understanding of inference in visual cortex and diffusion models</title>
    <link>https://arxiv.org/abs/2607.15693</link>
    <guid isPermaLink="false">paperfeed:2026-W29:arxiv:2607.15693</guid>
<pubDate>Sun, 30 Aug 2026 20:51:58 +0000</pubDate>    <description>Builds a recurrent sparse-coding V1 model that is mathematically equivalent to a minimal diffusion model, trained by denoising score matching, whose learned lateral interactions mirror V1 horizontal connections and whose Jacobian exposes how global consistency is enforced. A rare, interpretable mechanistic bridge between cortical circuits and diffusion inference. Look for: The cognitive/mechanistic claims outrun the quantitative evidence in the abstract—check how closely the learned connectivity matches real V1 data and the strength of the denoising comparison to black-box diffusion. (Zeyu Yun, Alexander Belsten, Dasheng Bi et al., https://arxiv.org/abs/2607.15693)</description>
  </item>
  <item>
    <title>9. Towards Human-level Dexterous Teleoperation</title>
    <link>https://arxiv.org/abs/2607.11481</link>
    <guid isPermaLink="false">paperfeed:2026-W29:arxiv:2607.11481</guid>
<pubDate>Sun, 30 Aug 2026 20:51:58 +0000</pubDate>    <description>TeleDexter jointly co-tracks hand and manipulated object with learned low-level contact behaviors, achieving zero-shot transfer to two real dexterous hands across seven reorientation and tool-use tasks at 75% success where baselines fail, and the demos train autonomous BC policies. A concrete real-world dexterity jump (grasp changes, in-hand manipulation, tool use). Look for: Baseline details and task difficulty are thin; assess whether &#39;all baselines fail&#39; reflects weak baselines and how demanding the seven tasks really are. (Puhao Li, Zeyuan Chen, Yingying Wu et al., https://arxiv.org/abs/2607.11481)</description>
  </item>
  <item>
    <title>10. Verbalizable Representations Form a Global Workspace in Language Models</title>
    <link>https://arxiv.org/abs/2607.15495</link>
    <guid isPermaLink="false">paperfeed:2026-W29:arxiv:2607.15495</guid>
<pubDate>Sun, 30 Aug 2026 20:51:58 +0000</pubDate>    <description>Introduces a Jacobian-based interpretability method identifying representations a language model is &#39;poised to verbalize&#39; (the J-space) and argues these exhibit global-workspace properties: reportable, held in memory, broadcast widely, active only in mid layers. A genuinely new framing for mechanistic interpretability that also surfaces hidden deliberation. Look for: The consciousness/global-workspace analogy is doing heavy lifting; read critically whether the J-space is a real functional bottleneck or a suggestive correlational construct. (Wes Gurnee, Nicholas Sofroniew, Adam Pearce et al., https://arxiv.org/abs/2607.15495)</description>
  </item>
  <item>
    <title>Issue 28 · Jul 6–12, 2026</title>
    <link>https://paperfeed.app/2026-W28/index.html</link>
    <guid isPermaLink="true">https://paperfeed.app/2026-W28/index.html</guid>
<pubDate>Sun, 30 Aug 2026 19:52:19 +0000</pubDate>    <description>This week is unusually rich in assumption-breaking results: a linear model for odor-mixture perception, a simple geometric baseline that beats deep nets for cross-session EEG, and a hidden gauge bug in a repetition penalty shipped across every major inference stack. On the capability side, a brain-to-voice BCI cuts word error 8x toward conversational quality, and code agents reportedly close full formal-verification coverage. We lean toward papers that overturn a common belief or demonstrate a real jump, with breadth across neuroscience, BCI, robotics, and LLM systems; several splashy world-model and video papers are held in mentions pending harder evidence. • 1. Odors Smell Like Their Components: A Linear Framework for Predicting Olfactory Mixture Perception • 2. Brain2voice 2.0: High-performance voice synthesis brain-computer interface • 3. Harnessing Code Agents for Automatic Software Verification • 4. Gauge dependence and structured-output corruption in sign-branched repetition penalties: measurements across models, inference stacks, and alternative repetition controls • 5. Constrained Decoding for Diffusion Language Models via Efficient Inference over Finite Automata • 6. Simple Geometric Recentering Rivals Deep Sequence Models for Cross-Session EEG Motor-Imagery Decoding • 7. Multiplayer Interactive World Models with Representation Autoencoders • 8. Computational demands shape seizure susceptibility in recurrent neural networks • 9. Disturbance-aware Motion Planning for Over-actuated Underwater Vehicles Exploiting Actuation Redundancy for High-fidelity 3D Reconstruction • 10. A Function-Space Dichotomy for Compositional Learning: Exponential Sub-Optimality of the Neural Tangent Kernel</description>
  </item>
  <item>
    <title>1. Odors Smell Like Their Components: A Linear Framework for Predicting Olfactory Mixture Perception</title>
    <link>https://paperfeed.app/2026-W28/biorxiv-10-64898-2026-07-03-736426.html</link>
    <guid isPermaLink="false">paperfeed:2026-W28:biorxiv:10.64898/2026.07.03.736426</guid>
<pubDate>Sun, 30 Aug 2026 19:52:19 +0000</pubDate>    <description>Directly overturns the entrenched assumption that odor-mixture perception is dominated by nonlinear receptor/neural interactions, showing simple averaging of component profiles predicts 432 mixtures near the noise ceiling. If it holds, it makes computational &#39;odorimetry&#39; tractable much like colorimetry, a genuine reframing of olfactory coding. Look for: Check whether the trained-panel quality descriptors and the linear model&#39;s success on previously &#39;emergent&#39; mixtures generalize beyond the specific odorant panel and concentration regime. (Pellegrino, R., Mayhew, E. J., Margolis, J. et al., https://www.biorxiv.org/content/10.64898/2026.07.03.736426v1)</description>
  </item>
  <item>
    <title>2. Brain2voice 2.0: High-performance voice synthesis brain-computer interface</title>
    <link>https://paperfeed.app/2026-W28/biorxiv-10-64898-2026-06-30-735633.html</link>
    <guid isPermaLink="false">paperfeed:2026-W28:biorxiv:10.64898/2026.06.30.735633</guid>
<pubDate>Sun, 30 Aug 2026 19:52:19 +0000</pubDate>    <description>An 8x reduction in word error (43.75% to 5.24%) for real-time intracortical voice synthesis is a major capability jump that pushes neural speech restoration toward practical conversational use. The causal 10ms multimodal decoder combining phoneme and acoustic targets is a concrete architecture, not just a benchmark bump. Look for: Note that evaluation is on a prior benchmark dataset with unspecified participant breadth; watch for generalization across speakers and truly online (not offline-rescored) performance. (Wairagkar, M., Srinivasan, A., Card, N. S. et al., https://www.biorxiv.org/content/10.64898/2026.06.30.735633v1)</description>
  </item>
  <item>
    <title>3. Harnessing Code Agents for Automatic Software Verification</title>
    <link>https://paperfeed.app/2026-W28/arxiv-2607-06341.html</link>
    <guid isPermaLink="false">paperfeed:2026-W28:arxiv:2607.06341</guid>
<pubDate>Sun, 30 Aug 2026 19:52:19 +0000</pubDate>    <description>Claims that handing whole lemmas to a general code agent with a verification harness beats fixed retrieval/tactic pipelines and achieves full coverage across 4,257 Iris lemmas and two proof assistants, a jump from ~1/8 coverage. If reproducible, this substantially resets expectations for LLM-driven formal verification. Look for: Scrutinize cost, per-lemma compute, harness engineering effort, and whether &#39;every lemma proved&#39; survives independent reproduction rather than curated targets. (Shuangxiang Kan, Shuanglong Kan, Sebastian Ertel, https://arxiv.org/abs/2607.06341)</description>
  </item>
  <item>
    <title>4. Gauge dependence and structured-output corruption in sign-branched repetition penalties: measurements across models, inference stacks, and alternative repetition controls</title>
    <link>https://paperfeed.app/2026-W28/arxiv-2607-09791.html</link>
    <guid isPermaLink="false">paperfeed:2026-W28:arxiv:2607.09791</guid>
<pubDate>Sun, 30 Aug 2026 19:52:19 +0000</pubDate>    <description>Identifies a widespread, previously hidden inference bug: the multiplicative repetition penalty branches on an arbitrary logit zero-point, so re-centering (a softmax no-op) changes 58-96% of greedy tokens and drops valid JSON from 97% to 23%. This affects HuggingFace, vLLM, llama.cpp and has a simple principled fix. Look for: Confirm the effect sizes replicate on larger RLHF checkpoints and that the normalized-logprob alternative doesn&#39;t introduce its own quality regressions. (Peter Hollows, https://arxiv.org/abs/2607.09791)</description>
  </item>
  <item>
    <title>5. Constrained Decoding for Diffusion Language Models via Efficient Inference over Finite Automata</title>
    <link>https://paperfeed.app/2026-W28/arxiv-2607-07026.html</link>
    <guid isPermaLink="false">paperfeed:2026-W28:arxiv:2607.07026</guid>
<pubDate>Sun, 30 Aug 2026 19:52:19 +0000</pubDate>    <description>A genuinely tailored solution to a central obstacle for diffusion LMs: exact finite-automaton constrained decoding despite parallel multi-token updates, with logarithmic-depth inference and large gains (22.3% to 69.0% on BFCL-Live) at under 5% overhead. This is a real capability enabler, not a decoding tweak. Look for: Watch how well the arithmetic-circuit depth reduction holds up in wall-clock terms across constraint complexity and whether gains persist beyond the two tested diffusion models. (Meihua Dang, Stefano Ermon, https://arxiv.org/abs/2607.07026)</description>
  </item>
  <item>
    <title>6. Simple Geometric Recentering Rivals Deep Sequence Models for Cross-Session EEG Motor-Imagery Decoding</title>
    <link>https://paperfeed.app/2026-W28/biorxiv-10-64898-2026-07-07-736991.html</link>
    <guid isPermaLink="false">paperfeed:2026-W28:biorxiv:10.64898/2026.07.07.736991</guid>
<pubDate>Sun, 30 Aug 2026 19:52:19 +0000</pubDate>    <description>A controlled eight-dataset benchmark showing a compact tangent-space classifier with unsupervised test-time recentering decisively beats deep Mamba-based decoders cross-session, with recentering (not model capacity) as the key factor. This challenges the field&#39;s drift toward ever-more-complex EEG architectures. Look for: The claim hinges on identical covariance features; check that the deep baselines were fairly tuned and that the within- vs cross-session dissociation is the true mechanism. (Rahimipour, M., Van Hulle, M., https://www.biorxiv.org/content/10.64898/2026.07.07.736991v1)</description>
  </item>
  <item>
    <title>7. Multiplayer Interactive World Models with Representation Autoencoders</title>
    <link>https://paperfeed.app/2026-W28/arxiv-2607-05352.html</link>
    <guid isPermaLink="false">paperfeed:2026-W28:arxiv:2607.05352</guid>
<pubDate>Sun, 30 Aug 2026 19:52:19 +0000</pubDate>    <description>A 5B latent-diffusion world model that conditions on all four players&#39; action streams in a fast, tightly coupled game, staying coherent far beyond its short training horizon at real-time framerates. Explicitly modeling multiple interacting action streams is a genuinely new world-model direction with released code and data. Look for: Be skeptical of the long-horizon stability and &#39;physical understanding&#39; claims; look for the quantitative distributional-quality metrics rather than the anecdotal hours-long rollouts. (Anthony Hu, Václav Volhejn, Adrien Ramanana Rahary et al., https://arxiv.org/abs/2607.05352)</description>
  </item>
  <item>
    <title>8. Computational demands shape seizure susceptibility in recurrent neural networks</title>
    <link>https://paperfeed.app/2026-W28/biorxiv-10-64898-2026-07-02-735135.html</link>
    <guid isPermaLink="false">paperfeed:2026-W28:biorxiv:10.64898/2026.07.02.735135</guid>
<pubDate>Sun, 30 Aug 2026 19:52:19 +0000</pubDate>    <description>Connects computation to pathology: RNN models predict that continuous-attractor representations are more seizure-vulnerable than discrete-state ones, and in vivo entorhinal vs CA3 recordings plus a connectivity manipulation support it. A non-obvious computational principle for regional seizure susceptibility that bridges modeling and neuroscience. Look for: Evidence is limited to a small set of regions and seizure conditions; check how robustly the attractor-type distinction maps onto the recorded dynamics. (Li, M., Eydam, S., Ramzan, I. et al., https://www.biorxiv.org/content/10.64898/2026.07.02.735135v1)</description>
  </item>
  <item>
    <title>9. Disturbance-aware Motion Planning for Over-actuated Underwater Vehicles Exploiting Actuation Redundancy for High-fidelity 3D Reconstruction</title>
    <link>https://paperfeed.app/2026-W28/arxiv-2607-07139.html</link>
    <guid isPermaLink="false">paperfeed:2026-W28:arxiv:2607.07139</guid>
<pubDate>Sun, 30 Aug 2026 19:52:19 +0000</pubDate>    <description>Demonstrates a broadly useful actuation-to-perception principle: exploiting an over-actuated ROV&#39;s thruster null space to minimize self-induced turbulence near the imaging target, cutting particle velocity 67% and reconstruction RMSE from 4.3mm to 1.9mm across 440 trials. Using control redundancy to improve sensing is a transferable idea beyond underwater robotics. Look for: Assess how dependent the wake proxy and allocator are on this specific eight-thruster platform and whether the principle transfers to other over-actuated systems. (Yuer Gao, Tongqing Xu, Qingyang Liu et al., https://arxiv.org/abs/2607.07139)</description>
  </item>
  <item>
    <title>10. A Function-Space Dichotomy for Compositional Learning: Exponential Sub-Optimality of the Neural Tangent Kernel</title>
    <link>https://paperfeed.app/2026-W28/arxiv-2607-06382.html</link>
    <guid isPermaLink="false">paperfeed:2026-W28:arxiv:2607.06382</guid>
<pubDate>Sun, 30 Aug 2026 19:52:19 +0000</pubDate>    <description>Gives a quantitative, function-space account of why finite trained ReLU nets beat their NTK limit on compositional targets, proving an exponential sample-complexity gap (4^L vs polynomial) with matching experiments on sparse parity. This is the kind of &#39;why it works&#39; result that explains a known failure mode rather than nudging a benchmark. Look for: The clean separation is on structured targets (iterated sawtooth, parity); consider how much the dichotomy informs realistic architectures and data beyond the unit circle/Boolean-cube settings. (Arkaprabha Ganguli, Emil Constantinescu, https://arxiv.org/abs/2607.06382)</description>
  </item>
  <item>
    <title>Project 1: EPFL-VILAB/Modus (GitHub)</title>
    <link>https://paperfeed.app/2026-W28/github-EPFL-VILAB-Modus.html</link>
    <guid isPermaLink="false">paperfeed:2026-W28:github:EPFL-VILAB/Modus</guid>
<pubDate>Sun, 30 Aug 2026 19:52:19 +0000</pubDate>    <description>A single decoder-only causal transformer that generates symmetrically across 16 modalities (text, images, depth, segmentation, detection, learned representations) without modality-specific heads is a genuinely new architecture direction, not a multimodal fine-tune. Full open release of training code, weights, data, and a demo, with ICML acceptance, makes it examinable rather than a promise. Look for: Check whether per-modality generation quality actually competes with specialist models or whether the unification comes at a steep quality cost. (https://github.com/EPFL-VILAB/Modus)</description>
  </item>
  <item>
    <title>Project 2: Robbyant/lingbot-world-v2 (GitHub)</title>
    <link>https://paperfeed.app/2026-W28/github-Robbyant-lingbot-world-v2.html</link>
    <guid isPermaLink="false">paperfeed:2026-W28:github:Robbyant/lingbot-world-v2</guid>
<pubDate>Sun, 30 Aug 2026 19:52:19 +0000</pubDate>    <description>A 14B interactive world model with causal generation for effectively unbounded rollouts, a distilled real-time 720p/60fps variant, and an agentic director harness is the strongest of this week&#39;s several world-model releases. Inference code and weights are out, and 1,500+ stars in days signals the field is treating it as a milestone. Look for: Verify the real-time distilled variant is actually released and how quickly scene coherence degrades over long interaction horizons. (https://github.com/Robbyant/lingbot-world-v2)</description>
  </item>
  <item>
    <title>Project 3: Helldez/BigMoeOnEdge (GitHub)</title>
    <link>https://paperfeed.app/2026-W28/github-Helldez-BigMoeOnEdge.html</link>
    <guid isPermaLink="false">paperfeed:2026-W28:github:Helldez/BigMoeOnEdge</guid>
<pubDate>Sun, 30 Aug 2026 19:52:19 +0000</pubDate>    <description>Lossless, byte-identical CPU inference of 284B-class MoE models on a 12 GB phone by streaming selected experts from flash is a substantial jump on the efficiency dimension, not a quantization trick. It works on stock llama.cpp, which makes it immediately reproducible and likely to be widely adopted. Look for: Check real tokens/sec on your target hardware and flash-wear implications, since expert-streaming throughput depends heavily on storage read bandwidth and routing locality. (https://github.com/Helldez/BigMoeOnEdge)</description>
  </item>
  <item>
    <title>Project 4: MuyeHuang/DuplexOmni (GitHub)</title>
    <link>https://paperfeed.app/2026-W28/github-MuyeHuang-DuplexOmni.html</link>
    <guid isPermaLink="false">paperfeed:2026-W28:github:MuyeHuang/DuplexOmni</guid>
<pubDate>Sun, 30 Aug 2026 19:52:19 +0000</pubDate>    <description>Full-duplex streaming audio+video in, speech out, with a fast interaction layer and a pluggable slower System-2 reasoning layer, targets exactly the gap between offline omni models and real interactive agents. The release is unusually complete: data generation, training, modified-vLLM serving, weights, and reproducible metadata. Look for: Test actual barge-in latency and turn-taking behavior yourself; full-duplex demos often degrade badly outside curated conditions. (https://github.com/MuyeHuang/DuplexOmni)</description>
  </item>
  <item>
    <title>Project 5: Robbyant/lingbot-video (GitHub)</title>
    <link>https://paperfeed.app/2026-W28/github-Robbyant-lingbot-video.html</link>
    <guid isPermaLink="false">paperfeed:2026-W28:github:Robbyant/lingbot-video</guid>
<pubDate>Sun, 30 Aug 2026 19:52:19 +0000</pubDate>    <description>A 30B MoE (3B active) video generator explicitly pretrained on 70k+ hours of embodied data with reward signals for physical plausibility and task completion is a serious open contribution to video-as-world-model for robotics. The full stack — models, prompt rewriters, inference code, and an eval benchmark — is released. Look for: Probe physical-plausibility quality against Wan/Cosmos baselines, since the README doesn&#39;t establish how strong the physics actually is. (https://github.com/Robbyant/lingbot-video)</description>
  </item>
  <item>
    <title>Project 6: WeZZard/jlens-qwen36 (GitHub)</title>
    <link>https://paperfeed.app/2026-W28/github-WeZZard-jlens-qwen36.html</link>
    <guid isPermaLink="false">paperfeed:2026-W28:github:WeZZard/jlens-qwen36</guid>
<pubDate>Sun, 30 Aug 2026 19:52:19 +0000</pubDate>    <description>This makes Jacobian-lens interpretability plus causal latent-state editing — including backward search for edits that produce a desired output — runnable locally on a consumer Mac, the most practical of this week&#39;s cluster of J-lens tools. Custom Metal kernels and released lens weights show real implementation depth beyond a paper port. Look for: The bundled lens is demo-grade and specific to Qwen3.6-27B; check how faithfully lens readouts track behavior before drawing scientific conclusions. (https://github.com/WeZZard/jlens-qwen36)</description>
  </item>
  <item>
    <title>Issue 27 · Jun 29 – Jul 5, 2026</title>
    <link>https://paperfeed.app/2026-W27/index.html</link>
    <guid isPermaLink="true">https://paperfeed.app/2026-W27/index.html</guid>
<pubDate>Sun, 30 Aug 2026 08:38:26 +0000</pubDate>    <description>This week leans heavy on assumption-breaking results and explanations of why things work: an LLM-discovered, machine-verified quantum proof; evidence that LLM &#39;evolution&#39; adds nothing over independent sampling; a token-level account of scaling laws; and a surprising claim that reliability scales inversely with model size. Neuroscience and BCI are unusually strong, headlined by near-implant-level non-invasive brain-to-text and a replicated multiregional Alzheimer&#39;s atlas. Robotics and speech round out the breadth, though many of the flashier system claims rest on single-benchmark or single-platform evidence and deserve scrutiny. • 1. A Machine-Verified Proof of a Quantum-Optimization Conjecture • 2. Accurate Decoding of Natural Sentences from Non-Invasive Brain Recordings • 3. Dictionaries, Not Darwin: Set-Level Selection Beats LLM Evolution in Scientific Equation Discovery • 4. Smooth Scaling Laws Hide Stepwise Token Learning • 5. Multiregional single-cell profiling reveals shared and specialized cellular vulnerability in Alzheimer&#39;s disease • 6. DeepGaze3.5-VL: Modeling Scanpaths via Autoregressive Token Prediction • 7. Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models • 8. Scaling Storm-Resolving Atmospheric AI Simulation to the Entire Planet • 9. Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents • 10. MorphQuad: Morphable Quadrotor for Superhuman Maneuverability, Manipulation, and Resiliency</description>
  </item>
  <item>
    <title>1. A Machine-Verified Proof of a Quantum-Optimization Conjecture</title>
    <link>https://paperfeed.app/2026-W27/arxiv-2606-29687.html</link>
    <guid isPermaLink="false">paperfeed:2026-W27:arxiv:2606.29687</guid>
<pubDate>Sun, 30 Aug 2026 08:38:26 +0000</pubDate>    <description>A decade-old QAOA conjecture resolved by an LLM that discovered a hidden dynamical symmetry, with the full proof mechanically checked in Lean 4. This is a rare, concrete demonstration of an LLM producing nontrivial new mathematical structure rather than reproducing known results, and the machine verification largely removes the usual trust problem. Look for: Check that the Lean formalization of the FGG statement faithfully encodes the intended conjecture, since the human-verified part is exactly the scaffolding. (Uri Kol, Maor Ben-Shahar, Kfir Sulimany et al., https://arxiv.org/abs/2606.29687)</description>
  </item>
  <item>
    <title>2. Accurate Decoding of Natural Sentences from Non-Invasive Brain Recordings</title>
    <link>https://paperfeed.app/2026-W27/arxiv-2608-18114.html</link>
    <guid isPermaLink="false">paperfeed:2026-W27:arxiv:2608.18114</guid>
<pubDate>Sun, 30 Aug 2026 08:38:26 +0000</pubDate>    <description>Non-invasive MEG decoding of naturally typed sentences reaching 39% WER, with the best subject getting half of sentences within one word error and log-linear improvement with data. This pushes non-invasive brain-to-text toward territory previously thought to require implants, which is directly in the reader&#39;s BCI wheelhouse. Look for: Note the small nine-subject cohort and likely heavy subject-specific training; the data-scaling extrapolation is the key claim to interrogate. (Mingfang Zhang, Jarod Lévy, Cedric Rommel et al., https://arxiv.org/abs/2608.18114)</description>
  </item>
  <item>
    <title>3. Dictionaries, Not Darwin: Set-Level Selection Beats LLM Evolution in Scientific Equation Discovery</title>
    <link>https://paperfeed.app/2026-W27/arxiv-2607-04108.html</link>
    <guid isPermaLink="false">paperfeed:2026-W27:arxiv:2607.04108</guid>
<pubDate>Sun, 30 Aug 2026 08:38:26 +0000</pubDate>    <description>Audits the popular &#39;LLM as evolutionary engine&#39; loop and finds parent-conditioned evolution is statistically indistinguishable from fresh independent sampling, then replaces it with a one-shot dictionary plus set-level sparse selection that beats the best baseline by a wide margin at a tenth the budget. A clean assumption-breaking result about a fast-growing methodology. Look for: Evidence is concentrated in scientific equation discovery; watch whether the &#39;evolution doesn&#39;t compound&#39; claim would hold in domains with reliable per-step credit. (Pan Li, https://arxiv.org/abs/2607.04108)</description>
  </item>
  <item>
    <title>4. Smooth Scaling Laws Hide Stepwise Token Learning</title>
    <link>https://paperfeed.app/2026-W27/arxiv-2606-29858.html</link>
    <guid isPermaLink="false">paperfeed:2026-W27:arxiv:2606.29858</guid>
<pubDate>Sun, 30 Aug 2026 08:38:26 +0000</pubDate>    <description>Offers a token-level explanation of why aggregate loss follows power laws: many contextualized tokens undergo sharp sigmoid learning transitions at different times, and their distribution reconstructs loss scaling across training, data, and model size. Backed by 100+ runs up to 6B/300B and an actionable 11% speedup from reweighting. Look for: Ask whether the sigmoid decomposition is causal/mechanistic or just an unusually good descriptive fit. (Pingjie Wang, Zechen Hu, Peiru Yang et al., https://arxiv.org/abs/2606.29858)</description>
  </item>
  <item>
    <title>5. Multiregional single-cell profiling reveals shared and specialized cellular vulnerability in Alzheimer&#39;s disease</title>
    <link>https://paperfeed.app/2026-W27/biorxiv-10-64898-2026-07-01-734821.html</link>
    <guid isPermaLink="false">paperfeed:2026-W27:biorxiv:10.64898/2026.07.01.734821</guid>
<pubDate>Sun, 30 Aug 2026 08:38:26 +0000</pubDate>    <description>A ~7M-nucleus, ten-region, 84-donor Alzheimer&#39;s atlas (replicated in 700+ donors) showing that only ~30% of cell types shift in abundance but do so coherently across regions, and that a supposedly resilient V1 layer-4 population becomes vulnerable. Large, replicated, and challenges assumptions about regional resilience. Look for: Mechanistic hyperexcitability conclusions remain hypothesis-generating; treat the vulnerability signatures as correlational. (Travaglini, K. J., Gabitto, M. I., Ding, Y. et al., https://www.biorxiv.org/content/10.64898/2026.07.01.734821v1)</description>
  </item>
  <item>
    <title>6. DeepGaze3.5-VL: Modeling Scanpaths via Autoregressive Token Prediction</title>
    <link>https://paperfeed.app/2026-W27/arxiv-2607-02083.html</link>
    <guid isPermaLink="false">paperfeed:2026-W27:arxiv:2607.02083</guid>
<pubDate>Sun, 30 Aug 2026 08:38:26 +0000</pubDate>    <description>Reframes human scanpath prediction as autoregressive token generation on a VLM, yielding a 46% information-gain jump over DeepGaze III that survives matched encoders, plus flexible conditioning and in-silico interventions that recover known oculomotor effects. A clean bridge between sequence modeling and perception that will interest both the AI and neuroscience sides. Look for: Verify the gain persists under identical encoders as claimed, and how much conditioning flexibility actually improves fit versus just adding capacity. (Susmit Agrawal, Matthias Bethge, Matthias Kümmerer, https://arxiv.org/abs/2607.02083)</description>
  </item>
  <item>
    <title>7. Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models</title>
    <link>https://paperfeed.app/2026-W27/arxiv-2607-18292.html</link>
    <guid isPermaLink="false">paperfeed:2026-W27:arxiv:2607.18292</guid>
<pubDate>Sun, 30 Aug 2026 08:38:26 +0000</pubDate>    <description>Claims reliability scales inversely: bigger models close the initial knowledge gap but degrade worse mid-response, driven by a per-token decoding-risk term that is invisible to the model&#39;s own uncertainty and grows with scale. If it holds, it&#39;s a genuinely surprising, self-perpetuating failure mode with a concrete mitigation. Look for: The causal language is very strong; scrutinize the oracle-based δ decomposition and whether the intervention truly targets risk rather than a proxy. (Kushal Chakrabarti, https://arxiv.org/abs/2607.18292)</description>
  </item>
  <item>
    <title>8. Scaling Storm-Resolving Atmospheric AI Simulation to the Entire Planet</title>
    <link>https://paperfeed.app/2026-W27/arxiv-2606-31248.html</link>
    <guid isPermaLink="false">paperfeed:2026-W27:arxiv:2606.31248</guid>
<pubDate>Sun, 30 Aug 2026 08:38:26 +0000</pubDate>    <description>First autoregressive AI emulator for global storm-resolving (~5 km) atmospheric dynamics, trained on tiles from just 17 days of data and blended into stable 24-hour global rollouts at ~50x the energy efficiency of the physics model. A substantive new direction in kilometer-scale climate/weather emulation. Look for: Only 17 training days and 24-hour horizons with acknowledged large-scale bias accumulation; the efficiency claim matters more than current fidelity. (Zeyuan Hu, Akshay Subramaniam, Noel Keen et al., https://arxiv.org/abs/2606.31248)</description>
  </item>
  <item>
    <title>9. Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents</title>
    <link>https://paperfeed.app/2026-W27/arxiv-2511-07397.html</link>
    <guid isPermaLink="false">paperfeed:2026-W27:arxiv:2511.07397</guid>
<pubDate>Sun, 30 Aug 2026 08:38:26 +0000</pubDate>    <description>Conversational infill lets a small &#39;talker&#39; model start answering immediately and fold in streamed reasoning/retrieval from a slower model, keeping millisecond time-to-first-response while closing much of the accuracy gap. A genuinely relevant architecture for real-time voice agents on the latency-capability frontier. Look for: Training relies on synthetic data, the accuracy-gap definition is fuzzy, and the user study is small (n=18). (Vidya Srinivas, Zachary Englhardt, Vikram Iyer et al., https://arxiv.org/abs/2511.07397)</description>
  </item>
  <item>
    <title>10. MorphQuad: Morphable Quadrotor for Superhuman Maneuverability, Manipulation, and Resiliency</title>
    <link>https://paperfeed.app/2026-W27/arxiv-2607-02764.html</link>
    <guid isPermaLink="false">paperfeed:2026-W27:arxiv:2607.02764</guid>
<pubDate>Sun, 30 Aug 2026 08:38:26 +0000</pubDate>    <description>A quadrotor with four independently gimbaled rotor assemblies plus a globally stable controller achieves omnidirectional flight, forceful contact tasks (turning valves, perching), and disturbance rejection in one compact platform. A real hardware-control co-design that expands what aerial manipulators can physically do. Look for: Few quantitative comparisons back the &#39;superhuman&#39; framing; look for hard numbers on force application and disturbance rejection. (Jose Diaz Peon Gonzalez Pacheco, Jiawei Xu, Andrew Zhao et al., https://arxiv.org/abs/2607.02764)</description>
  </item>
  <item>
    <title>Project 1: mira-wm/mira (GitHub)</title>
    <link>https://paperfeed.app/2026-W27/github-mira-wm-mira.html</link>
    <guid isPermaLink="false">paperfeed:2026-W27:github:mira-wm/mira</guid>
<pubDate>Sun, 30 Aug 2026 08:38:26 +0000</pubDate>    <description>Real-time interactive video world modeling with four concurrently controlled players at 20 FPS on a single GPU is a genuine capability jump, not a variant. The full release of training code, checkpoints, and a large multimodal game dataset makes it reproducible and buildable-upon. Look for: Check rollout coherence over long matches and whether the 20 FPS single-GPU claim holds outside cherry-picked clips or high-end hardware. (https://github.com/mira-wm/mira)</description>
  </item>
  <item>
    <title>Project 2: OpenSenseNova/SenseNova-Vision (GitHub)</title>
    <link>https://paperfeed.app/2026-W27/github-OpenSenseNova-SenseNova-Vision.html</link>
    <guid isPermaLink="false">paperfeed:2026-W27:github:OpenSenseNova/SenseNova-Vision</guid>
<pubDate>Sun, 30 Aug 2026 08:38:26 +0000</pubDate>    <description>Casting segmentation, depth, normals, and multi-view geometry as unified text/image generation without task-specific heads is a substantive architectural bet, backed by released 7B weights and a 50M-example corpus. If the formulation holds up, it points toward genuinely general vision models definable in natural language. Look for: Verify per-task performance against strong specialist baselines — unified formulations often trade accuracy on dense prediction tasks for generality. (https://github.com/OpenSenseNova/SenseNova-Vision)</description>
  </item>
  <item>
    <title>Project 3: meituan-longcat/LongCat-2.0 (GitHub)</title>
    <link>https://paperfeed.app/2026-W27/github-meituan-longcat-LongCat-2-0.html</link>
    <guid isPermaLink="false">paperfeed:2026-W27:github:meituan-longcat/LongCat-2.0</guid>
<pubDate>Sun, 30 Aug 2026 08:38:26 +0000</pubDate>    <description>A 1.6T-parameter open MoE trained on 35T tokens with a hardware-aware sparse attention design for million-token context is a frontier-scale release worth knowing regardless of benchmarks. The demonstrated large-scale training on non-GPU ASIC superpods is itself a notable data point about the hardware landscape. Look for: Check license terms, independent evaluations versus other open frontier models, and whether the 1M-context sparse attention actually delivers usable quality at that length. (https://github.com/meituan-longcat/LongCat-2.0)</description>
  </item>
  <item>
    <title>Project 4: anthropics/jacobian-lens (GitHub)</title>
    <link>https://paperfeed.app/2026-W27/github-anthropics-jacobian-lens.html</link>
    <guid isPermaLink="false">paperfeed:2026-W27:github:anthropics/jacobian-lens</guid>
<pubDate>Sun, 30 Aug 2026 08:38:26 +0000</pubDate>    <description>Transporting intermediate activations into the final-layer basis via averaged Jacobians is a principled interpretability method that exposes when internal representations become verbalizable — a global-workspace-flavored result connecting to the reader&#39;s neuroscience-AI interests. It ships as usable tooling with fitting support for open-weight models, not just paper code. Look for: Test whether the linear Jacobian-averaging approximation is faithful on the models you care about, or an artifact of the transport itself. (https://github.com/anthropics/jacobian-lens)</description>
  </item>
  <item>
    <title>Project 5: open-gigaai/giga-world-1 (GitHub)</title>
    <link>https://paperfeed.app/2026-W27/github-open-gigaai-giga-world-1.html</link>
    <guid isPermaLink="false">paperfeed:2026-W27:github:open-gigaai/giga-world-1</guid>
<pubDate>Sun, 30 Aug 2026 08:38:26 +0000</pubDate>    <description>An unusually complete open release for using video world models to evaluate robot policies — checkpoints, training recipes, data prep, and inference all included. Policy evaluation without real-robot rollouts is a bottleneck problem, and this is the most usable open stack for it this week. Look for: Note that distilled models, RL components, and parts of WMBench are unreleased, so validate whether world-model rollout scores actually correlate with real policy performance. (https://github.com/open-gigaai/giga-world-1)</description>
  </item>
  <item>
    <title>Project 6: XXH333/WordVoice-main (GitHub)</title>
    <link>https://paperfeed.app/2026-W27/github-XXH333-WordVoice-main.html</link>
    <guid isPermaLink="false">paperfeed:2026-W27:github:XXH333/WordVoice-main</guid>
<pubDate>Sun, 30 Aug 2026 08:38:26 +0000</pubDate>    <description>Explicit word-level planning and independent control of duration, energy, pitch, and intonation is the kind of decoupled prosody control TTS has lacked, and the release is unusually complete with training code, weights, and annotations. Directly relevant to the reader&#39;s speech interests as an editable-prosody direction. Look for: Listen to the demo and test whether the controls are actually independent and natural-sounding compared to existing controllable TTS, since no comparative evidence is provided. (https://github.com/XXH333/WordVoice-main)</description>
  </item>
  <item>
    <title>Issue 26 · Jun 22–28, 2026</title>
    <link>https://paperfeed.app/2026-W26/index.html</link>
    <guid isPermaLink="true">https://paperfeed.app/2026-W26/index.html</guid>
<pubDate>Sun, 30 Aug 2026 08:09:48 +0000</pubDate>    <description>This week&#39;s standout is a genuine landmark in computational imaging: the first complete virtual unrolling and scholarly reading of a sealed Herculaneum scroll. Robotics is strong on two fronts—sim-to-real methods that improve real policies without real-world data, and hardware/physics work that put a robot on the table against professional table-tennis players. Neuroscience delivers several assumption-challenging causal results (proprioception vs. reaching, spike synchrony vs. rate codes), and there&#39;s a cluster of omni-modal/audio-visual architectures plus theory explaining why RLVR beats SFT. • 1. Complete virtual unwrapping and reading of a rolled Herculaneum papyrus • 2. Support-Constrained RL Enables Real-World Policy Improvement without Real-World Experience • 3. Physics Models for Sim-to-Real Transfer in Professional-Level Robot Table Tennis • 4. Cervical spinal cord stimulation disrupts proprioception yet improves voluntary arm reaching • 5. The Importance of Synchrony in the Neural Control of Movement • 6. MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning • 7. Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models • 8. Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently • 9. A number simplex in the human medial temporal lobe • 10. Nemotron-Labs-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context</description>
  </item>
  <item>
    <title>1. Complete virtual unwrapping and reading of a rolled Herculaneum papyrus</title>
    <link>https://paperfeed.app/2026-W26/arxiv-2606-29085.html</link>
    <guid isPermaLink="false">paperfeed:2026-W26:arxiv:2606.29085</guid>
<pubDate>Sun, 30 Aug 2026 08:09:48 +0000</pubDate>    <description>A complete, papyrologically-reviewed virtual reading of an intact sealed scroll—moving from isolated patches to full unrolling, plus directly visible ink in another scroll and a title/author attribution. This is a demonstrated, potentially scalable capability that could unlock an entire ancient library. Look for: Check the coverage/review criteria and how much of the workflow is automated versus bespoke per-scroll, which determines whether it truly scales. (Giorgio Angelotti, Stephen Parsons, Federica Nicolardi et al., https://arxiv.org/abs/2606.29085)</description>
  </item>
  <item>
    <title>2. Support-Constrained RL Enables Real-World Policy Improvement without Real-World Experience</title>
    <link>https://paperfeed.app/2026-W26/arxiv-2606-27475.html</link>
    <guid isPermaLink="false">paperfeed:2026-W26:arxiv:2606.27475</guid>
<pubDate>Sun, 30 Aug 2026 08:09:48 +0000</pubDate>    <description>SCORE constrains simulated RL to the support of a real-data generative policy via flow steering, lifting eight real dexterous tasks from 37.8% to 89.9% with no additional real-world experience or distillation. It targets the core sim-to-real exploitation failure mode with a large, multi-task hardware result. Look for: Whether the support constraint limits improvement on tasks the base policy performs poorly, and how sensitive results are to the quality of the real-data prior. (Raymond Yu, William Huey, Mustafa Mukadam et al., https://arxiv.org/abs/2606.27475)</description>
  </item>
  <item>
    <title>3. Physics Models for Sim-to-Real Transfer in Professional-Level Robot Table Tennis</title>
    <link>https://paperfeed.app/2026-W26/arxiv-2606-28805.html</link>
    <guid isPermaLink="false">paperfeed:2026-W26:arxiv:2606.28805</guid>
<pubDate>Sun, 30 Aug 2026 08:09:48 +0000</pubDate>    <description>Broad, empirically calibrated flight and contact physics (drag/Magnus, table buckling, learned racket residuals) cut median landing error 59% and enable RL policies that compete with professional players. A rare case of a robot reaching a genuinely hard real-world skill via faithful modeling rather than benchmark gains. Look for: Details on opponent skill, match conditions, and win rates—the abstract asserts professional-level play without quantifying it. (Christian Conti, Bilan Yang, Alexander Sigrist et al., https://arxiv.org/abs/2606.28805)</description>
  </item>
  <item>
    <title>4. Cervical spinal cord stimulation disrupts proprioception yet improves voluntary arm reaching</title>
    <link>https://paperfeed.app/2026-W26/biorxiv-10-64898-2026-06-20-733548.html</link>
    <guid isPermaLink="false">paperfeed:2026-W26:biorxiv:10.64898/2026.06.20.733548</guid>
<pubDate>Sun, 30 Aug 2026 08:09:48 +0000</pubDate>    <description>A clean causal test: cervical spinal cord stimulation disrupts proprioception and postural stabilization yet improves rapid reaching smoothness and accuracy, helping resolve a decades-old debate about whether proprioception is required for goal-directed movement. Directly relevant to motor control and BCI/stimulation systems. Look for: Sample size and effect sizes are absent from the abstract; scrutinize how many participants and how robust the dissociation is. (Carranza, E., de Freitas, R., Verma, N. et al., https://www.biorxiv.org/content/10.64898/2026.06.20.733548v1)</description>
  </item>
  <item>
    <title>5. The Importance of Synchrony in the Neural Control of Movement</title>
    <link>https://paperfeed.app/2026-W26/biorxiv-10-64898-2026-06-26-734805.html</link>
    <guid isPermaLink="false">paperfeed:2026-W26:biorxiv:10.64898/2026.06.26.734805</guid>
<pubDate>Sun, 30 Aug 2026 08:09:48 +0000</pubDate>    <description>Millisecond-precise holographic optogenetics shows that motor-cortex output depends strongly on inter-neuron synchrony even when firing rates and cell identities are held fixed—causal evidence favoring a timing code over a rate code. This could change how we think about population coding of movement. Look for: Limited behavioral scope and quantitative detail; watch whether the synchrony dependence generalizes beyond the stimulation paradigm. (Hasegawa, M., Gruszka, B., Finch, M. S. et al., https://www.biorxiv.org/content/10.64898/2026.06.26.734805v1)</description>
  </item>
  <item>
    <title>6. MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning</title>
    <link>https://paperfeed.app/2026-W26/arxiv-2606-25225.html</link>
    <guid isPermaLink="false">paperfeed:2026-W26:arxiv:2606.25225</guid>
<pubDate>Sun, 30 Aug 2026 08:09:48 +0000</pubDate>    <description>A single unified encoder trained with one JEPA objective across audio and video, where cross-modal prediction is shown to be necessary for the shared representation to beat unimodal baselines. Strong frozen-evaluation gains with 10x less data make it a compelling alternative to modality-specific contrastive/reconstruction pipelines. Look for: Breadth of downstream evaluation—gains are concentrated on audio benchmarks, so check whether video representations hold up equally. (Revant Teotia, Adrien Bardes, Michael Rabbat et al., https://arxiv.org/abs/2606.25225)</description>
  </item>
  <item>
    <title>7. Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models</title>
    <link>https://paperfeed.app/2026-W26/arxiv-2606-25041.html</link>
    <guid isPermaLink="false">paperfeed:2026-W26:arxiv:2606.25041</guid>
<pubDate>Sun, 30 Aug 2026 08:09:48 +0000</pubDate>    <description>A single Transformer natively streaming interleaved language, audio, and video for full-duplex interaction, replacing cascaded ASR/LLM/TTS/avatar pipelines, with ~200 ms model latency. A genuinely ambitious any-to-any real-time architecture in a direction the reader tracks closely. Look for: Latency is well specified but capability/quality evidence is thin; be skeptical about response quality and how it compares to cascaded systems. (Lianghua Huang, Zhi-Fan Wu, Wei Wang et al., https://arxiv.org/abs/2606.25041)</description>
  </item>
  <item>
    <title>8. Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently</title>
    <link>https://paperfeed.app/2026-W26/arxiv-2606-22938.html</link>
    <guid isPermaLink="false">paperfeed:2026-W26:arxiv:2606.22938</guid>
<pubDate>Sun, 30 Aug 2026 08:09:48 +0000</pubDate>    <description>A theoretical account of why RLVR outperforms SFT: modeling CoT as graph pathfinding, SFT on shortest paths cannot learn to backtrack from dead ends, while outcome-reward RL can, yielding an exponential inference-compute separation. A crisp explanation for a live empirical debate. Look for: How restrictive the graph/pathfinding assumptions are and whether the exponential separation reflects realistic reasoning tasks. (Stanley Wei, Juno Kim, https://arxiv.org/abs/2606.22938)</description>
  </item>
  <item>
    <title>9. A number simplex in the human medial temporal lobe</title>
    <link>https://paperfeed.app/2026-W26/biorxiv-10-64898-2026-06-25-734462.html</link>
    <guid isPermaLink="false">paperfeed:2026-W26:biorxiv:10.64898/2026.06.25.734462</guid>
<pubDate>Sun, 30 Aug 2026 08:09:48 +0000</pubDate>    <description>Human medial-temporal-lobe recordings show number representations form high-dimensional simplex manifolds rather than a 1D mental number line, with linearly transferable codes across formats and a direct parallel to LLM representations plus attention-like arithmetic. A rare brain–AI computational connection. Look for: Robustness across subjects and tasks, and whether the LLM analogy is more than superficial geometric resemblance. (Zhu, H., Chericoni, A., Ismail, T. et al., https://www.biorxiv.org/content/10.64898/2026.06.25.734462v1)</description>
  </item>
  <item>
    <title>10. Nemotron-Labs-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context</title>
    <link>https://paperfeed.app/2026-W26/arxiv-2606-26493.html</link>
    <guid isPermaLink="false">paperfeed:2026-W26:arxiv:2606.26493</guid>
<pubDate>Sun, 30 Aug 2026 08:09:48 +0000</pubDate>    <description>TwoTower decouples context modeling (frozen AR tower) from diffusion denoising, retaining 98.7% of AR quality at 2.42x higher throughput on an open 30B hybrid MoE, with weights released. A clean architectural idea plus an open large-scale artifact make it more than a routine diffusion-LM tweak. Look for: Missing benchmark and hardware detail—verify the throughput claim reflects realistic serving rather than a favorable setting. (Fitsum Reda, John Kamalu, Roger Waleffe et al., https://arxiv.org/abs/2606.26493)</description>
  </item>
  <item>
    <title>Project 1: nicklashansen/mmbench2 (GitHub)</title>
    <link>https://paperfeed.app/2026-W26/github-nicklashansen-mmbench2.html</link>
    <guid isPermaLink="false">paperfeed:2026-W26:github:nicklashansen/mmbench2</guid>
<pubDate>Sun, 30 Aug 2026 08:09:48 +0000</pubDate>    <description>Hallucination is the central practical failure mode of learned world models, and this work claims it is both predictable and preventable, backed by an unusually complete release: 350M-param checkpoints across 210 tasks, 427 hours of video-action data, predictors, mitigation methods, and an interactive demo. This directly bears on whether world-model-based robot planning can be trusted, a live debate. Look for: Check whether the hallucination predictors generalize beyond the 350M scale and the ten benchmark domains, and whether the mitigation holds in closed-loop control rather than just open-loop rollouts. (https://github.com/nicklashansen/mmbench2)</description>
  </item>
  <item>
    <title>Project 2: vysri/conversational-infill (GitHub)</title>
    <link>https://paperfeed.app/2026-W26/github-vysri-conversational-infill.html</link>
    <guid isPermaLink="false">paperfeed:2026-W26:github:vysri/conversational-infill</guid>
<pubDate>Sun, 30 Aug 2026 08:09:48 +0000</pubDate>    <description>Framing the latency-accuracy tradeoff in voice agents as a trainable &#39;conversational infill&#39; task — a small local Talker speaks immediately while a cloud Reasoner streams knowledge into the ongoing response — is a genuinely new architecture for real-time voice systems. The release is complete: 290k-example dataset, seven Talker models, training code, and a live demo. Look for: Test how gracefully the Talker handles cases where the Reasoner&#39;s answer contradicts what it already said aloud, and whether infill quality holds outside the training distribution. (https://github.com/vysri/conversational-infill)</description>
  </item>
  <item>
    <title>Project 3: QwenLM/Qwen-AgentWorld (GitHub)</title>
    <link>https://paperfeed.app/2026-W26/github-QwenLM-Qwen-AgentWorld.html</link>
    <guid isPermaLink="false">paperfeed:2026-W26:github:QwenLM/Qwen-AgentWorld</guid>
<pubDate>Sun, 30 Aug 2026 08:09:48 +0000</pubDate>    <description>Training environment simulation as a native objective from continued pretraining — so one language model can simulate MCP tools, terminals, Android, web, and OS environments — is a substantive new direction for agent training, not an agent wrapper. Released MoE weights, trajectories, and a seven-domain benchmark make it testable, and it enables controllable and OOD simulation for RL without real environments. Look for: Frontier-level agent performance claims are self-reported; verify simulation fidelity on out-of-distribution tool behaviors before using it as a training environment. (https://github.com/QwenLM/Qwen-AgentWorld)</description>
  </item>
  <item>
    <title>Project 4: lil-lab/sps (GitHub)</title>
    <link>https://paperfeed.app/2026-W26/github-lil-lab-sps.html</link>
    <guid isPermaLink="false">paperfeed:2026-W26:github:lil-lab/sps</guid>
<pubDate>Sun, 30 Aug 2026 08:09:48 +0000</pubDate>    <description>Separating persistent state storage from next-token prediction via interleaved persistent and ephemeral streams in the KV cache is a genuinely different information-flow architecture, not an attention tweak. Reported gains hold from 53M to 1.68B params at matched inference cost, with reproducible training and ablation code from a credible lab. Look for: Check whether the scaling trend holds past 1.68B and whether the gains survive on downstream tasks rather than mainly language-modeling loss. (https://github.com/lil-lab/sps)</description>
  </item>
  <item>
    <title>Project 5: DAGroup-PKU/PhysisForcing (GitHub)</title>
    <link>https://paperfeed.app/2026-W26/github-DAGroup-PKU-PhysisForcing.html</link>
    <guid isPermaLink="false">paperfeed:2026-W26:github:DAGroup-PKU/PhysisForcing</guid>
<pubDate>Sun, 30 Aug 2026 08:09:48 +0000</pubDate>    <description>A training-time plug-in that reinforces physical plausibility in robotic video world models via interaction-region trajectories and relational constraints on DiT features, with zero inference overhead and released weights for Wan and Cosmos. The jump from 16% to 24% closed-loop planner success on WorldArena is a meaningful real-task result, not just video-quality metrics. Look for: Training code is not yet released and the closed-loop numbers are self-reported; verify the improvement replicates on planners other than the one evaluated. (https://github.com/DAGroup-PKU/PhysisForcing)</description>
  </item>
  <item>
    <title>Project 6: 0xShug0/audio.cpp (GitHub)</title>
    <link>https://paperfeed.app/2026-W26/github-0xShug0-audio-cpp.html</link>
    <guid isPermaLink="false">paperfeed:2026-W26:github:0xShug0/audio.cpp</guid>
<pubDate>Sun, 30 Aug 2026 08:09:48 +0000</pubDate>    <description>A pure C++ ggml-based inference engine unifying TTS, STT, VAD, voice conversion, and music generation with GGUF support and multiple GPU backends fills the same niche llama.cpp filled for LLMs — and its 2k stars in a week suggest it will be widely adopted. For anyone deploying local audio models, this could become default infrastructure. Look for: Check which specific models are actually supported with verified output parity versus merely listed, and how quickly new architectures get ported. (https://github.com/0xShug0/audio.cpp)</description>
  </item>
  <item>
    <title>Issue 25 · Jun 15–21, 2026</title>
    <link>https://paperfeed.app/2026-W25/index.html</link>
    <guid isPermaLink="true">https://paperfeed.app/2026-W25/index.html</guid>
<pubDate>Sun, 30 Aug 2026 07:42:19 +0000</pubDate>    <description>This week&#39;s strongest signal is a cluster of assumption-breaking results: AI now out-persuades expert human debaters and canvassers, glia turn out to release noradrenaline, medical VLMs often ignore the image, and simple geometry in frozen encoders beats giant world models at spotting physics violations. Robotics contributes real capability jumps (five-ball juggling, human-video-beats-robot-data pretraining) and neuroscience/BCI bring surprising mechanisms and a focal spinal-stimulation method. Many high-upvote LLM efficiency papers are interesting but under-evidenced from abstracts alone; treat their headline numbers skeptically. • 1. AI systems out-persuade expert humans • 2. A glial source of noradrenaline shapes synaptic integration and motor adaptation • 3. Task-Error Residual Learning for Real-Robot Five-Ball Juggling • 4. GEOPHYS: The Geometry of Physical Plausibility • 5. Vision-language models for chest radiography do not always need the image • 6. SPIDER -- Stitched Power-spectra for Inferring Directed information flow from incomplete and asynchronous Experimental Recordings • 7. HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining • 8. Adaptive Charge Modulation Enables Focal, Selective Spinal Cord Stimulation • 9. Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning • 10. Functional segregation of body-brain signals in the area postrema</description>
  </item>
  <item>
    <title>1. AI systems out-persuade expert humans</title>
    <link>https://paperfeed.app/2026-W25/arxiv-2606-16475.html</link>
    <guid isPermaLink="false">paperfeed:2026-W25:arxiv:2606.16475</guid>
<pubDate>Sun, 30 Aug 2026 07:42:19 +0000</pubDate>    <description>Four preregistered experiments (~19k conversations) showing frontier AI out-persuades laypeople, tournament winners, professional canvassers, and championship debaters—and transfers to real donations—is unusually strong, consequential evidence rather than a benchmark score. It also isolates a mechanism: the edge largely comes from rapidly deploying more information, and vanishes under human speed/length constraints. Look for: Check how persuasion was measured and whether the human-speed &#39;tie&#39; result generalizes beyond the specific coaching setup. (Kobi Hackenburg, Caroline Wagner, Luke Hewitt et al., https://arxiv.org/abs/2606.16475)</description>
  </item>
  <item>
    <title>2. A glial source of noradrenaline shapes synaptic integration and motor adaptation</title>
    <link>https://paperfeed.app/2026-W25/biorxiv-10-64898-2026-06-15-732457.html</link>
    <guid isPermaLink="false">paperfeed:2026-W25:biorxiv:10.64898/2026.06.15.732457</guid>
<pubDate>Sun, 30 Aug 2026 07:42:19 +0000</pubDate>    <description>A direct challenge to the canonical view that noradrenaline comes only from long-range nuclei: Bergmann glia are shown to synthesize and release noradrenaline via VMAT2, shaping Purkinje synaptic integration and motor adaptation. If it holds, it reframes local neuromodulation. Look for: Preprint with limited effect sizes—scrutinize the specificity of the VMAT2/genetic manipulations and whether release is truly glial rather than contamination from sparse axons. (Mach, S., Royer, J., Niu, W. et al., https://www.biorxiv.org/content/10.64898/2026.06.15.732457v1)</description>
  </item>
  <item>
    <title>3. Task-Error Residual Learning for Real-Robot Five-Ball Juggling</title>
    <link>https://paperfeed.app/2026-W25/arxiv-2606-16978.html</link>
    <guid isPermaLink="false">paperfeed:2026-W25:arxiv:2606.16978</guid>
<pubDate>Sun, 30 Aug 2026 07:42:19 +0000</pubDate>    <description>Stable five-ball juggling on real anthropomorphic arms converging essentially after one failed attempt is a striking real-robot capability, paired with a concrete methodological finding that directional task-error feedback plus an informative prior are jointly necessary (a fixed-Jacobian Newton update wins). Look for: Note how much the idealized analytic prior and hardware calibration carry the result, and whether the approach extends beyond periodic, well-modeled tasks. (Kai Ploeger, Jan Peters, https://arxiv.org/abs/2606.16978)</description>
  </item>
  <item>
    <title>4. GEOPHYS: The Geometry of Physical Plausibility</title>
    <link>https://paperfeed.app/2026-W25/arxiv-2606-20707.html</link>
    <guid isPermaLink="false">paperfeed:2026-W25:arxiv:2606.20707</guid>
<pubDate>Sun, 30 Aug 2026 07:42:19 +0000</pubDate>    <description>Claims that five geometric properties of frozen image-encoder features detect physical implausibility far better and cheaper than V-JEPA 2, GPT-4o, Gemini, and a dozen video diffusion models—and correlate with human EEG to object-permanence violations. The combination of capability, efficiency, and a neural link is exactly the kind of &#39;why it works&#39; result worth reading. Look for: Verify that the near-perfect LikePhys/IntPhys2 numbers aren&#39;t exploiting dataset artifacts, and how the geometry signals behave outside curated physics benchmarks. (Christian Internò, Alexander Pondaven, Habon Issa et al., https://arxiv.org/abs/2606.20707)</description>
  </item>
  <item>
    <title>5. Vision-language models for chest radiography do not always need the image</title>
    <link>https://paperfeed.app/2026-W25/arxiv-2606-17710.html</link>
    <guid isPermaLink="false">paperfeed:2026-W25:arxiv:2606.17710</guid>
<pubDate>Sun, 30 Aug 2026 07:42:19 +0000</pubDate>    <description>An intervention-based audit shows a 119B multimodal chest-radiograph model is statistically indistinguishable from a 7B text-only baseline, and a text-only model matches radiologist accuracy while grounding at zero. This directly overturns the assumption that high medical-VLM accuracy demonstrates visual reasoning, and offers a reusable causal audit. Look for: Consider whether the finding is specific to finding-name-prior-heavy datasets and how the grounding metrics would transfer to other medical imaging tasks. (Mahshad Lotfinia, Sebastian Ziegelmayer, Lisa Adams et al., https://arxiv.org/abs/2606.17710)</description>
  </item>
  <item>
    <title>6. SPIDER -- Stitched Power-spectra for Inferring Directed information flow from incomplete and asynchronous Experimental Recordings</title>
    <link>https://paperfeed.app/2026-W25/arxiv-2606-22695.html</link>
    <guid isPermaLink="false">paperfeed:2026-W25:arxiv:2606.22695</guid>
<pubDate>Sun, 30 Aug 2026 07:42:19 +0000</pubDate>    <description>SPIDER tackles a genuine methodological gap—recovering directed, frequency-specific interactions from asynchronous, partially overlapping recordings with no shared clock—and validates across calcium imaging, Neuropixels, and human iEEG, revealing a theta-band feedforward hierarchy with hippocampal formation at its source across species. Look for: Scrutinize the consistency guarantees&#39; assumptions and whether the matrix-completion step for never-co-observed regions introduces spurious directed flow. (Yisi S. Zhang, Daniel Y. Takahashi, https://arxiv.org/abs/2606.22695)</description>
  </item>
  <item>
    <title>7. HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining</title>
    <link>https://paperfeed.app/2026-W25/arxiv-2606-20521.html</link>
    <guid isPermaLink="false">paperfeed:2026-W25:arxiv:2606.20521</guid>
<pubDate>Sun, 30 Aug 2026 07:42:19 +0000</pubDate>    <description>A controlled comparison finding that egocentric human video, under a good filtering pipeline, outperforms teleoperated robot data for embodied pretraining—24% lower action loss and 90% higher out-of-distribution success—reverses a widely held assumption about the best pretraining source. Look for: Check how much of the gain is due to the filtering/labeling pipeline versus the data source, and whether the data budgets are truly matched. (Juncheng Ma, Jianxin Bi, Yufan Deng et al., https://arxiv.org/abs/2606.20521)</description>
  </item>
  <item>
    <title>8. Adaptive Charge Modulation Enables Focal, Selective Spinal Cord Stimulation</title>
    <link>https://paperfeed.app/2026-W25/biorxiv-10-64898-2026-06-16-732691.html</link>
    <guid isPermaLink="false">paperfeed:2026-W25:biorxiv:10.64898/2026.06.16.732691</guid>
<pubDate>Sun, 30 Aug 2026 07:42:19 +0000</pubDate>    <description>Adaptive Charge Modulation achieves focal, deep spinal activation with single-muscle selectivity from epidural surface electrodes—normally requiring implanted contacts—with chronic 68-day stability and dense 2,112-channel brain-spine recordings. A genuinely new neuromodulation strategy relevant to neural interfaces. Look for: Evidence is rat-only; watch for how selectivity was quantified and whether the high-frequency suppression mechanism is established rather than inferred. (Vatsyayan, R., Khoury, F., Porter, T. S. et al., https://www.biorxiv.org/content/10.64898/2026.06.16.732691v1)</description>
  </item>
  <item>
    <title>9. Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning</title>
    <link>https://paperfeed.app/2026-W25/arxiv-2606-18134.html</link>
    <guid isPermaLink="false">paperfeed:2026-W25:arxiv:2606.18134</guid>
<pubDate>Sun, 30 Aug 2026 07:42:19 +0000</pubDate>    <description>Diarization-conditioning the acoustic encoder of a spoken LLM while keeping the decoder frozen yields large speaker-attributed transcription gains over Gemini 3 Flash and Voxtral on far-field multi-talker audio, plus strong long-form multi-speaker QA. A concrete architectural fix for a real speech-model limitation. Look for: Assess fairness of the baseline comparisons and whether gains depend on the specific DiCoW/Voxtral pairing versus the general conditioning idea. (Alexander Polok, Samuele Cornell, Sathvik Udupa et al., https://arxiv.org/abs/2606.18134)</description>
  </item>
  <item>
    <title>10. Functional segregation of body-brain signals in the area postrema</title>
    <link>https://paperfeed.app/2026-W25/biorxiv-10-64898-2026-06-15-732473.html</link>
    <guid isPermaLink="false">paperfeed:2026-W25:biorxiv:10.64898/2026.06.15.732473</guid>
<pubDate>Sun, 30 Aug 2026 07:42:19 +0000</pubDate>    <description>A functional remapping of area postrema cell types—GFRAL neurons (canonically sickness/nausea) responding to dietary fat independently of GDF15, with GIPR neurons gating them via sugar—connects widely used weight-loss-drug targets to natural nutrient sensing in a non-obvious way. Look for: Mouse physiology only; check the causal specificity of the GDF15-independent fat pathway and how cell-type identity was verified. (Lopez-Cruz, A., Burgos, N. S. F., Hakimi, A. M. et al., https://www.biorxiv.org/content/10.64898/2026.06.15.732473v1)</description>
  </item>
  <item>
    <title>Project 1: kingjulio8238/nanoG1 (GitHub)</title>
    <link>https://paperfeed.app/2026-W25/github-kingjulio8238-nanoG1.html</link>
    <guid isPermaLink="false">paperfeed:2026-W25:github:kingjulio8238/nanoG1</guid>
<pubDate>Sun, 30 Aug 2026 07:42:19 +0000</pubDate>    <description>Training a Unitree G1 walking policy from scratch in ~59 seconds on one GPU, via a robot-specialized compiled physics engine hitting 7.25M steps/s, is a genuine orders-of-magnitude efficiency jump for humanoid RL. It ships the full stack: simulation, training code, browser demo, and real-hardware deployment, making it both a capability result and a reusable tool. Look for: Verify the real-robot gait quality and robustness beyond flat-ground walking, and whether the specialized physics engine generalizes to other tasks or is locked to G1 locomotion. (https://github.com/kingjulio8238/nanoG1)</description>
  </item>
  <item>
    <title>Project 2: catnip-ai-tech/MaineCoon (GitHub)</title>
    <link>https://paperfeed.app/2026-W25/github-catnip-ai-tech-MaineCoon.html</link>
    <guid isPermaLink="false">paperfeed:2026-W25:github:catnip-ai-tech/MaineCoon</guid>
<pubDate>Sun, 30 Aug 2026 07:42:19 +0000</pubDate>    <description>Treating synchronized audio-video generation as a native streaming problem — sub-second interaction and up to 47.5 FPS from a 22B model on a single H100 — is a meaningful shift from adapting offline diffusion, and directly in the reader&#39;s omni-modal/interactive-systems lane. If the claims hold, this is the kind of real-time social world model people will build on. Look for: The repo appears to be mostly a technical report and links; check whether weights or runnable code exist and whether the latency/FPS numbers are independently reproducible. (https://github.com/catnip-ai-tech/MaineCoon)</description>
  </item>
  <item>
    <title>Project 3: LogosRoboticsGroup/A2World (GitHub)</title>
    <link>https://paperfeed.app/2026-W25/github-LogosRoboticsGroup-A2World.html</link>
    <guid isPermaLink="false">paperfeed:2026-W25:github:LogosRoboticsGroup/A2World</guid>
<pubDate>Sun, 30 Aug 2026 07:42:19 +0000</pubDate>    <description>An action-conditioned multi-view diffusion world model pretrained on 2.1M manipulation trajectories across 20+ embodiments, with one dynamics prior transferring to both policy-evaluation simulation and instruction-conditioned control. Code and checkpoints are released, making it one of the more substantive robot world-model artifacts this year. Look for: Evidence is mostly from the authors&#39; own benchmarks with near-zero external traction; check rollout fidelity over long horizons and whether the released checkpoints reproduce the reported transfer results. (https://github.com/LogosRoboticsGroup/A2World)</description>
  </item>
  <item>
    <title>Project 4: facebookresearch/kernel_bench_verified (GitHub)</title>
    <link>https://paperfeed.app/2026-W25/github-facebookresearch-kernel-bench-verified.html</link>
    <guid isPermaLink="false">paperfeed:2026-W25:github:facebookresearch/kernel_bench_verified</guid>
<pubDate>Sun, 30 Aug 2026 07:42:19 +0000</pubDate>    <description>Shows that reported LLM-generated CUDA kernel speedups can collapse from 1.43x to 0.88x under realistic TF32 baselines and hidden correctness tests — a result that overturns a widely repeated claim about LLM kernel generation. This is exactly the &#39;explains why something works (or doesn&#39;t)&#39; finding the reader values, from Facebook Research. Look for: Very new with minimal adoption; verify the baseline choices are fair and see whether the original KernelBench authors or kernel-generation groups respond to the methodology. (https://github.com/facebookresearch/kernel_bench_verified)</description>
  </item>
  <item>
    <title>Project 5: ShareLab-SII/UniAR (GitHub)</title>
    <link>https://paperfeed.app/2026-W25/github-ShareLab-SII-UniAR.html</link>
    <guid isPermaLink="false">paperfeed:2026-W25:github:ShareLab-SII/UniAR</guid>
<pubDate>Sun, 30 Aug 2026 07:42:19 +0000</pubDate>    <description>The claim that the visual tokenizer, not the architecture, is the key to unifying understanding, generation, and editing in one autoregressive model is a clean, testable design insight (ICML 2026) — the model can read its own generated tokens without re-encoding. Code, checkpoints, and a live demo are all released. Look for: Visual-decoder training code is withheld and traction is modest; test the demo yourself on editing chains where re-encoding artifacts would normally accumulate. (https://github.com/ShareLab-SII/UniAR)</description>
  </item>
  <item>
    <title>Project 6: Multimedia-Semantic-Analytics-Lab/PerceptionDLM (GitHub)</title>
    <link>https://paperfeed.app/2026-W25/github-Multimedia-Semantic-Analytics-Lab-PerceptionDLM.html</link>
    <guid isPermaLink="false">paperfeed:2026-W25:github:Multimedia-Semantic-Analytics-Lab/PerceptionDLM</guid>
<pubDate>Sun, 30 Aug 2026 07:42:19 +0000</pubDate>    <description>Using diffusion-LM parallel denoising to caption many image regions simultaneously sidesteps autoregressive latency scaling with region count — a genuinely different decoding regime for perception with a reported 3.4x throughput gain. Full release of code, 8B weights, training data, and a benchmark makes it immediately usable. Look for: Check per-region caption quality against strong autoregressive VLMs at matched compute, since parallel decoding often trades fidelity for throughput. (https://github.com/Multimedia-Semantic-Analytics-Lab/PerceptionDLM)</description>
  </item>
  <item>
    <title>Issue 24 · Jun 8–14, 2026</title>
    <link>https://paperfeed.app/2026-W24/index.html</link>
    <guid isPermaLink="true">https://paperfeed.app/2026-W24/index.html</guid>
<pubDate>Sun, 30 Aug 2026 07:14:35 +0000</pubDate>    <description>A strong week for efficiency and new framings in transformers (editable/composable KV caches, group-wise sparse attention), plus several assumption-overturning neuroscience results and a genuinely novel BCI platform. Robotics contributes two data-efficiency jumps (human-video-to-dexterous-robot, zero-shot sim-to-real deformables). As always, discount the systems/benchmark overclaims where abstracts report only best-case numbers; the neuroscience picks are preprints, so treat mechanisms as provisional. • 1. Models Take Notes at Prefill: KV Cache Can Be Editable and Composable • 2. A Fully Endovascular Neural Interface • 3. Dexterous Point Policy: Learning Point-based Dexterous Hand Policies from Human Demonstrations • 4. The Signs Were Always There: Training-Free Concept Detection and Steering in Raw Transformer Dimensions • 5. MiniMax Sparse Attention • 6. Dynamic trajectory cues drive sequenced integration in approach detectors • 7. Lineage tracing and live-cell imaging reveal that NeuroD1 does not reprogram microglia into neurons • 8. A Two-Dimensional Grid-Cell Code for Three-Dimensional Navigation in Freely Flying Bats • 9. SimWeaver: Zero-Shot RGB Sim-to-Real for Deformable Manipulation • 10. Overcoming State Inertia in Full-Duplex Spoken Language Models via Activation Steering</description>
  </item>
  <item>
    <title>1. Models Take Notes at Prefill: KV Cache Can Be Editable and Composable</title>
    <link>https://paperfeed.app/2026-W24/arxiv-2606-17107.html</link>
    <guid isPermaLink="false">paperfeed:2026-W24:arxiv:2606.17107</guid>
<pubDate>Sun, 30 Aug 2026 07:14:35 +0000</pubDate>    <description>Reframes the KV cache as a notebook of memoized, field-conditioned conclusions that can be edited after a correction and RoPE-repositioned/spliced into new contexts, with causal evidence across four model families. If it holds up it is both a new conceptual lens on what prefill computes and a practical serving win (large TTFT reductions, append-only, composes with prefix caching). Look for: Check the causal claim that the field&#39;s own KV drives &lt;1% of the decision, and whether edit+compose stays decision-identical outside the curated benchmarks and for non-CoT settings. (Bojie Li, https://arxiv.org/abs/2606.17107)</description>
  </item>
  <item>
    <title>2. A Fully Endovascular Neural Interface</title>
    <link>https://paperfeed.app/2026-W24/biorxiv-10-64898-2026-06-07-730604.html</link>
    <guid isPermaLink="false">paperfeed:2026-W24:biorxiv:10.64898/2026.06.07.730604</guid>
<pubDate>Sun, 30 Aug 2026 07:14:35 +0000</pubDate>    <description>A fully endovascular, sub-1-mm3, ultrasound-powered neural implant delivered like a stent, demonstrating autonomic stimulation and blood-pressure modulation in rabbits. This is a genuinely new, less-invasive neural-interface platform rather than an incremental electrode improvement. Look for: Note that this is stimulation only in a small-animal acute setting; recording capability, chronic safety, and durability remain unproven. (Stanton, J., Talei Franzesi, G., Spinazzi, E. et al., https://www.biorxiv.org/content/10.64898/2026.06.07.730604v1)</description>
  </item>
  <item>
    <title>3. Dexterous Point Policy: Learning Point-based Dexterous Hand Policies from Human Demonstrations</title>
    <link>https://paperfeed.app/2026-W24/arxiv-2606-10614.html</link>
    <guid isPermaLink="false">paperfeed:2026-W24:arxiv:2606.10614</guid>
<pubDate>Sun, 30 Aug 2026 07:14:35 +0000</pubDate>    <description>Learns dexterous multi-finger manipulation from human videos with zero robot demonstrations by using a shared wrist+fingertip 3D keypoint representation for both observation and action, reporting 75% vs 1% for a VLA baseline. The claim that keypoint-level alignment largely dissolves the human-to-dexterous-robot embodiment gap would be a meaningful data-efficiency jump. Look for: The task suite and baseline breadth are thin in the abstract; scrutinize how many tasks, how the 1% VLA baseline was configured, and whether keypoints capture contact-rich forces. (Beomjun Kim, Seong Hyeon Park, Seunghoon Sim et al., https://arxiv.org/abs/2606.10614)</description>
  </item>
  <item>
    <title>4. The Signs Were Always There: Training-Free Concept Detection and Steering in Raw Transformer Dimensions</title>
    <link>https://paperfeed.app/2026-W24/arxiv-2606-12629.html</link>
    <guid isPermaLink="false">paperfeed:2026-W24:arxiv:2606.12629</guid>
<pubDate>Sun, 30 Aug 2026 07:14:35 +0000</pubDate>    <description>Argues the raw standard basis of transformer hidden states is a training-free, cross-modal feature basis: signs encode content, per-dim reading loses nothing over a full MLP, and flipping sign patterns causally steers concepts. If robust, this undercuts a core premise of SAE/dictionary interpretability and separates reader vs writer dimensions. Look for: Sweeping claims and only 2 upvotes—verify the steering results and that &#39;sign alone&#39; truly rivals learned probes rather than reflecting benchmark-specific artifacts. (Varun Reddy Nalagatla, https://arxiv.org/abs/2606.12629)</description>
  </item>
  <item>
    <title>5. MiniMax Sparse Attention</title>
    <link>https://paperfeed.app/2026-W24/arxiv-2606-13392.html</link>
    <guid isPermaLink="false">paperfeed:2026-W24:arxiv:2606.13392</guid>
<pubDate>Sun, 30 Aug 2026 07:14:35 +0000</pubDate>    <description>A streamlined block-sparse attention over GQA with group-specific top-k selection and a co-designed kernel, reporting 28.4x attention-compute reduction and 14.2x/7.6x prefill/decode speedups at 1M context on a 109B multimodal model. This is the kind of simple, deployable long-context method that tends to get widely adopted. Look for: Task-level quality parity at 1M tokens is under-reported; check retrieval/agentic quality, not just perplexity, and how it compares to other learned-sparse schemes. (Xunhao Lai, Weiqi Xu, Yufeng Yang et al., https://arxiv.org/abs/2606.13392)</description>
  </item>
  <item>
    <title>6. Dynamic trajectory cues drive sequenced integration in approach detectors</title>
    <link>https://paperfeed.app/2026-W24/biorxiv-10-64898-2026-06-05-730527.html</link>
    <guid isPermaLink="false">paperfeed:2026-W24:biorxiv:10.64898/2026.06.05.730527</guid>
<pubDate>Sun, 30 Aug 2026 07:14:35 +0000</pubDate>    <description>Shows luminance change alone evokes an approach/retreat percept in both humans and flies, identifies dual-purpose approach-detector neurons in Drosophila, and finds cues are integrated synergistically only in the natural temporal order. A clean cross-species link from a new percept to a defined, sequence-sensitive circuit computation. Look for: How strong the causal silencing/imaging evidence is for the &#39;sequenced integration&#39; claim versus a correlational temporal-order effect. (Vashistha, H., Matos, N. C., Wu, H. et al., https://www.biorxiv.org/content/10.64898/2026.06.05.730527v1)</description>
  </item>
  <item>
    <title>7. Lineage tracing and live-cell imaging reveal that NeuroD1 does not reprogram microglia into neurons</title>
    <link>https://paperfeed.app/2026-W24/biorxiv-10-64898-2026-06-08-730780.html</link>
    <guid isPermaLink="false">paperfeed:2026-W24:biorxiv:10.64898/2026.06.08.730780</guid>
<pubDate>Sun, 30 Aug 2026 07:14:35 +0000</pubDate>    <description>Using virus-free lineage tracing, longitudinal two-photon imaging, and scRNA-seq, finds NeuroD1 does not convert microglia to neurons and instead drives microglial apoptosis—directly challenging a prominent and contested glia-to-neuron reprogramming literature. Methodologically rigorous negative result that could recontextualize many prior conversion claims. Look for: The abstract&#39;s final sentence appears to contain a contradiction/typo; read the actual lineage-tracing controls to confirm the direction of the claim. (Li, X., Li, Y., Cao, Y. et al., https://www.biorxiv.org/content/10.64898/2026.06.08.730780v1)</description>
  </item>
  <item>
    <title>8. A Two-Dimensional Grid-Cell Code for Three-Dimensional Navigation in Freely Flying Bats</title>
    <link>https://paperfeed.app/2026-W24/biorxiv-10-64898-2026-06-10-728358.html</link>
    <guid isPermaLink="false">paperfeed:2026-W24:biorxiv:10.64898/2026.06.10.728358</guid>
<pubDate>Sun, 30 Aug 2026 07:14:35 +0000</pubDate>    <description>Wireless recordings from freely flying bats show grid cells retain a 2D toroidal manifold, and flight paths are organized along transient 2D planes—offering a concrete resolution to how a 2D grid code could support 3D navigation. From the Yartsev lab, this reframes a long-standing debate about grid coding in 3D. Look for: Whether the &#39;plane-of-motion&#39; account generalizes beyond structured foraging flights and how robustly the toroidal topology holds during genuinely volumetric maneuvers. (Qi, K. K., Yartsev, M. M., https://www.biorxiv.org/content/10.64898/2026.06.10.728358v1)</description>
  </item>
  <item>
    <title>9. SimWeaver: Zero-Shot RGB Sim-to-Real for Deformable Manipulation</title>
    <link>https://paperfeed.app/2026-W24/arxiv-2606-15338.html</link>
    <guid isPermaLink="false">paperfeed:2026-W24:arxiv:2606.15338</guid>
<pubDate>Sun, 30 Aug 2026 07:14:35 +0000</pubDate>    <description>Reports zero-shot RGB sim-to-real for visually complex deformable manipulation (plastic bags, silk) from 200 sim demos per task, with 91% average real success and strong robustness under visual shift where real-data baselines collapse. Deformable RGB sim-to-real without real fine-tuning has been largely unsolved. Look for: How much comes from the ISP-aware photometric augmentation and measurement-backed simulator; check whether success holds on objects far from the asset-generation distribution. (Wenkang Hu, Haoran Wang, Yitong Li et al., https://arxiv.org/abs/2606.15338)</description>
  </item>
  <item>
    <title>10. Overcoming State Inertia in Full-Duplex Spoken Language Models via Activation Steering</title>
    <link>https://paperfeed.app/2026-W24/arxiv-2606-11386.html</link>
    <guid isPermaLink="false">paperfeed:2026-W24:arxiv:2606.11386</guid>
<pubDate>Sun, 30 Aug 2026 07:14:35 +0000</pubDate>    <description>Diagnoses &#39;state inertia&#39; in full-duplex spoken LMs—internal representations stay biased toward generation just after a barge-in, causing the model to miss the start of user speech—and fixes it training-free via an activation-steering perception vector, with a zero-buffer benchmark and sizable gains. A crisp mechanistic insight into interactive speech models plus a practical intervention. Look for: Whether the steering vector generalizes across models and conversational conditions, and how latency/quality trade off in real deployment. (Cheng-Kuang Chang, Kai-Wei Chang, Alexander H. Liu et al., https://arxiv.org/abs/2606.11386)</description>
  </item>
  <item>
    <title>Project 1: jd-opensource/JoyAI-VL-Interaction (GitHub)</title>
    <link>https://paperfeed.app/2026-W24/github-jd-opensource-JoyAI-VL-Interaction.html</link>
    <guid isPermaLink="false">paperfeed:2026-W24:github:jd-opensource/JoyAI-VL-Interaction</guid>
<pubDate>Sun, 30 Aug 2026 07:14:35 +0000</pubDate>    <description>An open 8B system that continuously watches video and autonomously decides when to speak, stay silent, or delegate is exactly the proactive real-time multimodal direction the reader tracks. The release is unusually complete — model, training recipe, time-aligned interaction data, quantized checkpoints, and deployment stack — rather than an offline model with a demo video. Look for: Verify actual end-to-end latency and speak/silence decision quality on your own streams; check whether the interaction data license permits derivative training. (https://github.com/jd-opensource/JoyAI-VL-Interaction)</description>
  </item>
  <item>
    <title>Project 2: NVlabs/SpatialClaw (GitHub)</title>
    <link>https://paperfeed.app/2026-W24/github-NVlabs-SpatialClaw.html</link>
    <guid isPermaLink="false">paperfeed:2026-W24:github:NVlabs/SpatialClaw</guid>
<pubDate>Sun, 30 Aug 2026 07:14:35 +0000</pubDate>    <description>Replacing rigid tool-calling with a persistent Python kernel the VLM programs against — with segmentation, depth, and geometry tools whose intermediate results it can inspect — is a genuinely fresh action-interface idea. An 11-point average gain across 20 spatial benchmarks and six backbones, training-free, suggests it generalizes rather than overfitting one setup. Look for: Check inference cost per query (multi-step code execution can be slow/expensive) and whether the gains hold outside the curated benchmark suite. (https://github.com/NVlabs/SpatialClaw)</description>
  </item>
  <item>
    <title>Project 3: sbryngelson/ANEForge (GitHub)</title>
    <link>https://paperfeed.app/2026-W24/github-sbryngelson-ANEForge.html</link>
    <guid isPermaLink="false">paperfeed:2026-W24:github:sbryngelson/ANEForge</guid>
<pubDate>Sun, 30 Aug 2026 07:14:35 +0000</pubDate>    <description>Pure-ANE execution including on-engine backpropagation and Adam training from ordinary Python, bypassing CoreML entirely, is a capability nobody outside Apple has had. If the MLPerf-valid results hold, it materially changes what Apple-silicon developers can do with the previously opaque Neural Engine. Look for: It depends on private, unsupported APIs that could break with any macOS update — treat it as research infrastructure, not production, and verify the reported numbers on your own hardware. (https://github.com/sbryngelson/ANEForge)</description>
  </item>
  <item>
    <title>Project 4: allenai/molmo-motion (GitHub)</title>
    <link>https://paperfeed.app/2026-W24/github-allenai-molmo-motion.html</link>
    <guid isPermaLink="false">paperfeed:2026-W24:github:allenai/molmo-motion</guid>
<pubDate>Sun, 30 Aug 2026 07:14:35 +0000</pubDate>    <description>Language-conditioned 3D trajectory prediction for arbitrary user-selected points is a new intermediate representation between video models and robot policies, and the demonstrated transfer to both robot planning and motion-guided video generation is compelling. Allen AI releases the model, a million-example corpus, benchmarks, and training recipes. Look for: Check how well trajectory predictions hold up on cluttered real scenes versus curated evaluation data, and how much the robot-planning transfer depends on downstream machinery. (https://github.com/allenai/molmo-motion)</description>
  </item>
  <item>
    <title>Project 5: RightNow-AI/AutoMegaKernel (GitHub)</title>
    <link>https://paperfeed.app/2026-W24/github-RightNow-AI-AutoMegaKernel.html</link>
    <guid isPermaLink="false">paperfeed:2026-W24:github:RightNow-AI/AutoMegaKernel</guid>
<pubDate>Sun, 30 Aug 2026 07:14:35 +0000</pubDate>    <description>An agent harness that verifies, fuses, and self-tunes an entire Llama-style decode pass into one persistent CUDA megakernel — retargeting itself across GPU generations — is substantive automated systems engineering, not a wrapper. The honest reporting that its equal-precision bf16 path still loses to cuBLAS makes the int8 wins far more credible. Look for: Confirm the correctness gating covers your model variant and precision; gains are currently specific to batch-1 int8 decode on inference-class GPUs. (https://github.com/RightNow-AI/AutoMegaKernel)</description>
  </item>
  <item>
    <title>Project 6: HarryHsing/OmniAgent (GitHub)</title>
    <link>https://paperfeed.app/2026-W24/github-HarryHsing-OmniAgent.html</link>
    <guid isPermaLink="false">paperfeed:2026-W24:github:HarryHsing/OmniAgent</guid>
<pubDate>Sun, 30 Aug 2026 07:14:35 +0000</pubDate>    <description>A 7B agent that natively decides which frames, audio, or clips to fetch while reasoning — beating a 72B model with 73% fewer frames — is a strong data point that active perception, not brute-force context, is the path for long video understanding. The turn-level RL credit assignment for perception actions is a nice methodological contribution too. Look for: Traction is minimal and benchmark details are thin; verify the LVBench comparison setup and whether weights and the RL training code are actually released. (https://github.com/HarryHsing/OmniAgent)</description>
  </item>
  <item>
    <title>Issue 23 · Jun 1–7, 2026</title>
    <link>https://paperfeed.app/2026-W23/index.html</link>
    <guid isPermaLink="true">https://paperfeed.app/2026-W23/index.html</guid>
<pubDate>Sun, 30 Aug 2026 06:41:18 +0000</pubDate>    <description>This week leans heavily toward doing more with less: several papers internalize context into parameters or memory (Frames2LoRA, cartridges, adaptive video codecs) and toward assumption-breaking diagnostics of what our models and benchmarks actually measure. Formal-math agents took a visible jump, physics-constrained generative inference got a foundational correction, and there&#39;s a rich neuroscience crop touching directly on computation, coding, and chaos. Note that many of the flashiest claims (Cosmos 3, several agent self-improvement results) rest on thin quantitative evidence in the abstracts, so discount accordingly. • 1. Frames2LoRA: Parametric Video Internalization for Vision-Language Models • 2. Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities • 3. Predictable Mean-Field Chaos in Random Recurrent Neural Networks • 4. Goedel-Architect: Streamlining Formal Theorem Proving with Blueprint Generation and Refinement • 5. Wave Focusing in Metamaterials: Tactile Displays Beyond the Diffraction Limit • 6. The Right Measure for Physics-Constrained Generation: A Co-Area Correction for Posterior-Consistent PDE Inverse Problems • 7. The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models • 8. What Are We Actually Benchmarking in Robot Manipulation? • 9. dots.tts Technical Report • 10. Intrinsic Population Dynamics are a Neuronal Substrate for Visual Attention</description>
  </item>
  <item>
    <title>1. Frames2LoRA: Parametric Video Internalization for Vision-Language Models</title>
    <link>https://paperfeed.app/2026-W23/arxiv-2606-04351.html</link>
    <guid isPermaLink="false">paperfeed:2026-W23:arxiv:2606.04351</guid>
<pubDate>Sun, 30 Aug 2026 06:41:18 +0000</pubDate>    <description>A genuinely new framing: a hypernetwork reads a VLM&#39;s layerwise activations while it encodes a video and emits a LoRA adapter in one forward pass, so the video lives in weights rather than context. The reported 6–80x latency and up to 1,500x token reductions with non-inferior quality, plus rank-space composition of independently generated adapters, point at a real new direction for long-video and memory. Look for: Check whether &#39;statistical non-inferiority&#39; hides systematic quality loss on harder QA, and how well adapter composition actually holds beyond a few chunks. (Manan Suri, Sarvesh Baskar, Dinesh Manocha, https://arxiv.org/abs/2606.04351)</description>
  </item>
  <item>
    <title>2. Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities</title>
    <link>https://paperfeed.app/2026-W23/arxiv-2608-18090.html</link>
    <guid isPermaLink="false">paperfeed:2026-W23:arxiv:2608.18090</guid>
<pubDate>Sun, 30 Aug 2026 06:41:18 +0000</pubDate>    <description>Claims a single valence direction recoverable from just nine emotion anchors that appears in text, vision, audio, AND human EEG encoders never jointly trained, with causal ablation effects in LLMs. If the cross-modal alignment and controls hold, this is an unusually strong statement about shared representational geometry across models and brains. Look for: Scrutinize the EEG and cross-modal probe alignment for confounds, and note the honest caveats: bounded to continuous attributes and family-specific steering. (Yousef Radwan, https://arxiv.org/abs/2608.18090)</description>
  </item>
  <item>
    <title>3. Predictable Mean-Field Chaos in Random Recurrent Neural Networks</title>
    <link>https://paperfeed.app/2026-W23/arxiv-2606-08805.html</link>
    <guid isPermaLink="false">paperfeed:2026-W23:arxiv:2606.08805</guid>
<pubDate>Sun, 30 Aug 2026 06:41:18 +0000</pubDate>    <description>A striking theoretical result: a random recurrent network can have a positive Lyapunov exponent yet be perfectly predictable at the single-neuron level from its continuous past, showing predictive complexity and microscopic instability scale differently. This directly bears on how we interpret neural variability and chaos in both brains and RNNs. Look for: The result leans on continuous-time DMFT idealizations; look for finite-network, finite-sampling validation and how the log-p horizon degrades under noise. (Alkesh Yadav, Vladimir Shaidurov, Jonathan Kadmon, https://arxiv.org/abs/2606.08805)</description>
  </item>
  <item>
    <title>4. Goedel-Architect: Streamlining Formal Theorem Proving with Blueprint Generation and Refinement</title>
    <link>https://paperfeed.app/2026-W23/arxiv-2606-06468.html</link>
    <guid isPermaLink="false">paperfeed:2026-W23:arxiv:2606.06468</guid>
<pubDate>Sun, 30 Aug 2026 06:41:18 +0000</pubDate>    <description>Reframes theorem proving around a global dependency-graph blueprint that is refined on failure rather than recursively decomposed, reporting 99.2% MiniF2F, 75.6% PutnamBench, and solutions to recent olympiad problems at a claimed ~500x lower cost. This is a substantial capability-and-efficiency jump in formal math. Look for: Verify how much comes from the 284B backbone versus the blueprint method, and watch for benchmark contamination on recent competition problems. (Jui-Hui Chung, Ziyang Cai, Zihao Li et al., https://arxiv.org/abs/2606.06468)</description>
  </item>
  <item>
    <title>5. Wave Focusing in Metamaterials: Tactile Displays Beyond the Diffraction Limit</title>
    <link>https://paperfeed.app/2026-W23/arxiv-2606-05572.html</link>
    <guid isPermaLink="false">paperfeed:2026-W23:arxiv:2606.05572</guid>
<pubDate>Sun, 30 Aug 2026 06:41:18 +0000</pubDate>    <description>A real fabricated haptic display that uses a locally resonant metamaterial plate to focus tactile waves beyond the plate&#39;s diffraction limit, achieving a tenfold reduction in virtual-pixel area with few actuators and validated behaviorally. Genuinely new physics/hardware for distributed touch, backed by builds and human experiments rather than simulation. Look for: Note bandwidth/frequency constraints and whether independent multi-point control degrades as more simultaneous pixels are demanded. (Gregory Reardon, Max Linnander, Dustin Goetz et al., https://arxiv.org/abs/2606.05572)</description>
  </item>
  <item>
    <title>6. The Right Measure for Physics-Constrained Generation: A Co-Area Correction for Posterior-Consistent PDE Inverse Problems</title>
    <link>https://paperfeed.app/2026-W23/arxiv-2606-04804.html</link>
    <guid isPermaLink="false">paperfeed:2026-W23:arxiv:2606.04804</guid>
<pubDate>Sun, 30 Aug 2026 06:41:18 +0000</pubDate>    <description>Argues that the standard recipe of projecting a generative prior onto a hard PDE constraint samples the wrong posterior because it omits a co-area (Fixman) Jacobian, and shows the bias is large (up to 20x the noise floor). This is a potentially foundational correction for the fast-growing area of generative PDE inverse problems. Look for: Evidence is on controlled problems with an i.i.d. arbiter; watch whether CoCoS remains tractable and accurate on realistic high-dimensional inverse tasks. (Jian Xu, Yanning Wu, Delu Zeng et al., https://arxiv.org/abs/2606.04804)</description>
  </item>
  <item>
    <title>7. The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models</title>
    <link>https://paperfeed.app/2026-W23/arxiv-2606-05976.html</link>
    <guid isPermaLink="false">paperfeed:2026-W23:arxiv:2606.05976</guid>
<pubDate>Sun, 30 Aug 2026 06:41:18 +0000</pubDate>    <description>A clean, surprising result: LLMs&#39; failure to correct their own reasoning errors is largely an artifact of the chat template&#39;s role labeling, not a cognitive deficit—relabeling identical erroneous text as an external role raises correction rates by 23–93 points. This reframes self-correction and how we evaluate instruction tuning. Look for: Check robustness across chat templates and whether the fix generalizes beyond the tested math/logic tasks or just gates explicit flagging. (Kuan-Yen Chen, Fang-Yi Su, Shih-Yen Lin et al., https://arxiv.org/abs/2606.05976)</description>
  </item>
  <item>
    <title>8. What Are We Actually Benchmarking in Robot Manipulation?</title>
    <link>https://paperfeed.app/2026-W23/arxiv-2606-04233.html</link>
    <guid isPermaLink="false">paperfeed:2026-W23:arxiv:2606.04233</guid>
<pubDate>Sun, 30 Aug 2026 06:41:18 +0000</pubDate>    <description>Concrete audits showing popular manipulation benchmarks (LIBERO, CALVIN) are weak proxies for capability: a 0.09B language-free probe hits near-SOTA on LIBERO, most gains aren&#39;t statistically significant, and modest within-range pose randomization breaks CALVIN policies. Provides reusable diagnostics the field badly needs. Look for: Consider whether the four diagnostics themselves fully capture real-world manipulation ability, and how the audited benchmarks compare to RoboCasa/RoboTwin. (Tianchong Jiang, Xiangshan Tan, Samuel Wheeler et al., https://arxiv.org/abs/2606.04233)</description>
  </item>
  <item>
    <title>9. dots.tts Technical Report</title>
    <link>https://paperfeed.app/2026-W23/arxiv-2606-07080.html</link>
    <guid isPermaLink="false">paperfeed:2026-W23:arxiv:2606.07080</guid>
<pubDate>Sun, 30 Aug 2026 06:41:18 +0000</pubDate>    <description>A strong, fully open continuous-autoregressive TTS system with a prediction-friendly AudioVAE, full-history flow conditioning, reward-free self-correction, and 54–85 ms first-packet latency via MeanFlow distillation. Directly in the reader&#39;s speech/voice interest with credible quality and real-time claims plus released checkpoints. Look for: Comparisons are mostly open-source-SOTA framing; check multilingual robustness and how distillation affects expressiveness versus the non-distilled model. (Shi Lian, Changtao Li, Bohan Li et al., https://arxiv.org/abs/2606.07080)</description>
  </item>
  <item>
    <title>10. Intrinsic Population Dynamics are a Neuronal Substrate for Visual Attention</title>
    <link>https://paperfeed.app/2026-W23/biorxiv-10-64898-2026-06-02-729565.html</link>
    <guid isPermaLink="false">paperfeed:2026-W23:biorxiv:10.64898/2026.06.02.729565</guid>
<pubDate>Sun, 30 Aug 2026 06:41:18 +0000</pubDate>    <description>Reports structured, stimulus-independent &#39;blob-like&#39; population dynamics in the superior colliculus that emerge with learning, predict trial-by-trial behavior, and amplify sensory responses up to fourfold—casting intrinsic dynamics as an active attentional substrate rather than noise. A candidate shift in how we think about attention and intrinsic activity. Look for: Watch how strongly the causal claims are supported and whether the excitatory-inhibitory model is doing explanatory work or just fitting. (Schmidt, F. H., Mlynarski, W., Georges, A. et al., https://www.biorxiv.org/content/10.64898/2026.06.02.729565v1)</description>
  </item>
  <item>
    <title>Project 1: studio-dots-ai/dots.tts (GitHub)</title>
    <link>https://paperfeed.app/2026-W23/github-studio-dots-ai-dots-tts.html</link>
    <guid isPermaLink="false">paperfeed:2026-W23:github:studio-dots-ai/dots.tts</guid>
<pubDate>Sun, 30 Aug 2026 06:41:18 +0000</pubDate>    <description>A 2B fully-continuous autoregressive TTS with Apache-2.0 code and weights, strong multilingual zero-shot cloning, 48 kHz output, and streaming/distilled variants — this is exactly the kind of open speech release the reader tracks. Traction and completeness (training + inference code) suggest it will become a widely used baseline. Look for: Verify the multilingual and cloning benchmarks against Fish/CosyVoice-class systems independently, and check real streaming latency on your hardware rather than reported numbers. (https://github.com/studio-dots-ai/dots.tts)</description>
  </item>
  <item>
    <title>Project 2: jd-opensource/JoyAI-Echo (GitHub)</title>
    <link>https://paperfeed.app/2026-W23/github-jd-opensource-JoyAI-Echo.html</link>
    <guid isPermaLink="false">paperfeed:2026-W23:github:jd-opensource/JoyAI-Echo</guid>
<pubDate>Sun, 30 Aug 2026 06:41:18 +0000</pubDate>    <description>Maintaining cross-shot audiovisual continuity over ~5-minute videos plus an enterable, causally rolled-out world model with joint visual, ambient audio, music, and speech is a real capability step beyond short-clip video generation. Code, checkpoints, and strong early traction back it up. Look for: Check the actual quality and consistency of the 5-minute samples and the compute/inference profile needed — long-horizon claims often degrade badly past the cherry-picked demos. (https://github.com/jd-opensource/JoyAI-Echo)</description>
  </item>
  <item>
    <title>Project 3: facebookresearch/brain2qwerty (GitHub)</title>
    <link>https://paperfeed.app/2026-W23/github-facebookresearch-brain2qwerty.html</link>
    <guid isPermaLink="false">paperfeed:2026-W23:github:facebookresearch/brain2qwerty</guid>
<pubDate>Sun, 30 Aug 2026 06:41:18 +0000</pubDate>    <description>Sentence-level brain-to-text decoding from non-invasive MEG/EEG is a meaningful BCI capability advance, published in Nature Neuroscience with released code and a Spanish MEG/EEG dataset. Directly in the reader&#39;s BCI wheelhouse and reproducible enough to build on. Look for: Note that MEG requires a shielded room (not wearable), check the character error rates for EEG vs MEG separately, and that the v2 dataset remains embargoed. (https://github.com/facebookresearch/brain2qwerty)</description>
  </item>
  <item>
    <title>Project 4: akarshkumar0101/smt (GitHub)</title>
    <link>https://paperfeed.app/2026-W23/github-akarshkumar0101-smt.html</link>
    <guid isPermaLink="false">paperfeed:2026-W23:github:akarshkumar0101/smt</guid>
<pubDate>Sun, 30 Aug 2026 06:41:18 +0000</pubDate>    <description>Training nonlinear RNNs without backpropagating through recurrence — using a transformer teacher to generate one-step memory-transition targets — is a genuinely different framing that attacks BPTT&#39;s core parallelism and credit-assignment limits. Working PyTorch code and comparisons on language and pixel-sequence tasks make it more than a proposal. Look for: Check how it scales beyond small models and whether the DAgger-style correction handles compounding state drift at long horizons; the teacher-student dependency may cap final quality. (https://github.com/akarshkumar0101/smt)</description>
  </item>
  <item>
    <title>Project 5: 19PINE-AI/programmable-kv (GitHub)</title>
    <link>https://paperfeed.app/2026-W23/github-19PINE-AI-programmable-kv.html</link>
    <guid isPermaLink="false">paperfeed:2026-W23:github:19PINE-AI/programmable-kv</guid>
<pubDate>Sun, 30 Aug 2026 06:41:18 +0000</pubDate>    <description>Reframing the KV cache as editable, composable program state — with causal mechanism experiments showing prefill writes conclusions onto downstream tokens — is both an interpretability insight and a potentially important serving primitive (append-only errata, transplantable skills). Very novel, with code and multi-model results. Look for: It is days old with almost no external validation; test whether edits stay coherent under long generations and across model families before treating the claims as established. (https://github.com/19PINE-AI/programmable-kv)</description>
  </item>
  <item>
    <title>Project 6: divelab/OPDLM (GitHub)</title>
    <link>https://paperfeed.app/2026-W23/github-divelab-OPDLM.html</link>
    <guid isPermaLink="false">paperfeed:2026-W23:github:divelab/OPDLM</guid>
<pubDate>Sun, 30 Aug 2026 06:41:18 +0000</pubDate>    <description>On-policy distillation for converting pretrained AR language models into block-diffusion models, with released data, code, and 0.6B–8B checkpoints, is a substantive and directly testable contribution to the AR-vs-diffusion LM debate. Training the student on its own diffusion trajectories rather than teacher-forced targets is the interesting methodological piece. Look for: Compare the converted models&#39; quality/speed tradeoff against the AR originals and other diffusion LMs yourself — conversion papers often hide capability regressions on reasoning tasks. (https://github.com/divelab/OPDLM)</description>
  </item>
  <item>
    <title>Issue 22 · May 25–31, 2026</title>
    <link>https://paperfeed.app/2026-W22/index.html</link>
    <guid isPermaLink="true">https://paperfeed.app/2026-W22/index.html</guid>
<pubDate>Sun, 30 Aug 2026 06:15:12 +0000</pubDate>    <description>This week&#39;s strongest signals cluster around three themes: interpretability results that overturn assumptions (scaled mechanistic features in a production LLM, and evidence that probes and even fMRI foundation models measure the wrong thing), robotics design principles and simulators that unlock capabilities rather than nudge benchmarks, and new tooling for measuring computation in single neurons. Note that abstract dates are anomalously stamped in the future; several papers (e.g. C2) read like landmark work, so weigh claims against your own recollection. We kept the neuroscience and interpretability picks heavy because that&#39;s where the genuinely assumption-breaking evidence is this week. • 1. Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet • 2. Extreme dynamic symmetry enables omnidirectional and multifunctional robots • 3. Crazyflow: An Accurate, GPU-Accelerated, Differentiable Drone Simulator in JAX • 4. Ultrasensitive voltage imaging reveals distinct electrical microdomains in neurons • 5. The Variance Brain Foundation Models Forgot: Third-Order Statistics Predict Cognition Where Billion-Parameter Models Fail • 6. When and How Long? The Readout-Mediator Angle in Temporal Reasoning • 7. Learning to Search and Searching to Learn for Generalization in Planning • 8. Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players • 9. Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention • 10. When Does LeJEPA Learn a World Model?</description>
  </item>
  <item>
    <title>1. Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet</title>
    <link>https://paperfeed.app/2026-W22/arxiv-2605-29358.html</link>
    <guid isPermaLink="false">paperfeed:2026-W22:arxiv:2605.29358</guid>
<pubDate>Sun, 30 Aug 2026 06:15:12 +0000</pubDate>    <description>Scaling sparse autoencoders to Claude 3 Sonnet is the first strong evidence that dictionary-learning interpretability generalizes from toy transformers to production-scale, multimodal models, with 34M features and causal steering. This is the kind of result that shifts a whole subfield&#39;s expectations. Look for: The authors themselves flag that the feature set is incomplete and faithfulness is not rigorously established; check how much the causal steering demonstrations actually constrain the &#39;this is what the model computes&#39; interpretation. (Adly Templeton, Tom Conerly, Jonathan Marcus et al., https://arxiv.org/abs/2605.29358)</description>
  </item>
  <item>
    <title>2. Extreme dynamic symmetry enables omnidirectional and multifunctional robots</title>
    <link>https://paperfeed.app/2026-W22/arxiv-2605-29254.html</link>
    <guid isPermaLink="false">paperfeed:2026-W22:arxiv:2605.29254</guid>
<pubDate>Sun, 30 Aug 2026 06:15:12 +0000</pubDate>    <description>&#39;Dynamic symmetry&#39;—engineering a robot so attainable center-of-mass accelerations are isotropic—is a genuinely new organizing principle for robot design, not a geometric or control tweak. It is backed by 1000+ morphology simulations and a physical 20-leg spherical robot demonstrating orientation-invariant locomotion and failure tolerance. Look for: Watch whether the benefits are truly attributable to dynamic isotropy versus the specific radial-linear-actuator architecture, and how the principle would transfer to more conventional morphologies. (Jiaxun Liu, Boxi Xia, Boyuan Chen, https://arxiv.org/abs/2605.29254)</description>
  </item>
  <item>
    <title>3. Crazyflow: An Accurate, GPU-Accelerated, Differentiable Drone Simulator in JAX</title>
    <link>https://paperfeed.app/2026-W22/arxiv-2606-01478.html</link>
    <guid isPermaLink="false">paperfeed:2026-W22:arxiv:2606.01478</guid>
<pubDate>Sun, 30 Aug 2026 06:15:12 +0000</pubDate>    <description>A differentiable, GPU-accelerated JAX drone simulator that is &gt;10x faster per drone and scales to thousands of 4000-drone swarms, but the real headline is breaking the train-then-deploy paradigm: training a recovery policy from scratch in 0.38s while a physical drone is airborne. That in-execution learning demonstration is the surprising part. Look for: Scrutinize how much of the sub-centimeter tracking and in-flight learning depends on the Crazyflie&#39;s specific dynamics being easy to model; generality to contact-rich or higher-dimensional systems is unproven. (Martin Schuck, Marcel P. Rath, Yufei Hua et al., https://arxiv.org/abs/2606.01478)</description>
  </item>
  <item>
    <title>4. Ultrasensitive voltage imaging reveals distinct electrical microdomains in neurons</title>
    <link>https://paperfeed.app/2026-W22/biorxiv-10-64898-2026-05-27-728040.html</link>
    <guid isPermaLink="false">paperfeed:2026-W22:biorxiv:10.64898/2026.05.27.728040</guid>
<pubDate>Sun, 30 Aug 2026 06:15:12 +0000</pubDate>    <description>ASAP7y is a voltage indicator with subthreshold sensitivity that, combined with EM reconstruction across 717 Drosophila cell types, provides mechanistic evidence that single neurons perform spatially localized, parallel computations rather than acting as uniform integrators. This is both an enabling measurement tool and a substantive claim about neural computation. Look for: The electrical-microdomain claims lean partly on electrotonic modeling from morphology; note where the conclusions are direct measurements versus model inference. (Hao, Y. A., Jayne, L. L., Lee, S. et al., https://www.biorxiv.org/content/10.64898/2026.05.27.728040v1)</description>
  </item>
  <item>
    <title>5. The Variance Brain Foundation Models Forgot: Third-Order Statistics Predict Cognition Where Billion-Parameter Models Fail</title>
    <link>https://paperfeed.app/2026-W22/arxiv-2606-04010.html</link>
    <guid isPermaLink="false">paperfeed:2026-W22:arxiv:2606.04010</guid>
<pubDate>Sun, 30 Aug 2026 06:15:12 +0000</pubDate>    <description>A pointed, assumption-breaking result: fMRI foundation models predict cognition worse than plain functional connectivity, and worse as they scale, because pretraining preserves second-order covariance but destroys the third-order co-skewness that carries cognitive signal. A no-GPU linear pipeline beats billion-parameter models, and targeted finetuning closes the gap—implicating the objective, not the architecture. Look for: Effect sizes and the specific readout/eval protocol; the &#39;variance allocation&#39; story is compelling but rests on a particular cumulant analysis and a small set of datasets/parcellations. (Giovanni Marraffini, Gabriel Mahuas, Trinidad Borrell et al., https://arxiv.org/abs/2606.04010)</description>
  </item>
  <item>
    <title>6. When and How Long? The Readout-Mediator Angle in Temporal Reasoning</title>
    <link>https://paperfeed.app/2026-W22/arxiv-2605-29126.html</link>
    <guid isPermaLink="false">paperfeed:2026-W22:arxiv:2605.29126</guid>
<pubDate>Sun, 30 Aug 2026 06:15:12 +0000</pubDate>    <description>A clean, causally-supported demonstration that a linear probe can decode a feature nearly perfectly while being orthogonal to the subspace the model actually uses (found via DAS), replicated across scales and families. This directly undermines a common inference from probing studies and matters for anyone reading interpretability results. Look for: Generality beyond the calendar-date/duration task; the spatial and arithmetic extensions are described as preliminary. (Shreyas Fadnavis, Praitayini Kanakaraj, Felix Wyss, https://arxiv.org/abs/2605.29126)</description>
  </item>
  <item>
    <title>7. Learning to Search and Searching to Learn for Generalization in Planning</title>
    <link>https://paperfeed.app/2026-W22/arxiv-2605-25720.html</link>
    <guid isPermaLink="false">paperfeed:2026-W22:arxiv:2605.25720</guid>
<pubDate>Sun, 30 Aug 2026 06:15:12 +0000</pubDate>    <description>A search-to-learning loop pairing a relational GNN heuristic with weighted A* and Q-learning yields striking zero-shot combinatorial generalization—heuristics trained on &lt;30-block Blocksworld solving 488-block instances without search—across Sokoban, PushWorld, The Witness, and IPC. This is a real capability jump on generalization in planning. Look for: The abstract is thin on quantitative comparisons; check how brittle the 30-to-488 transfer is to problem structure and whether it holds outside relational domains with clean state descriptions. (Michael Aichmüller, Yannik Hesse, Hector Geffner, https://arxiv.org/abs/2605.25720)</description>
  </item>
  <item>
    <title>8. Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players</title>
    <link>https://paperfeed.app/2026-W22/arxiv-2605-28816.html</link>
    <guid isPermaLink="false">paperfeed:2026-W22:arxiv:2605.28816</guid>
<pubDate>Sun, 30 Aug 2026 06:15:12 +0000</pubDate>    <description>An interactive video world model for multiple independently controlled agents, with permutation-symmetric agent encodings, hub-mediated linear cross-agent attention, and real-time 24-FPS causal generation that generalizes from two to four players without retraining. Multi-agent interactive world models are an underexplored and genuinely new direction, and this is the most upvoted paper of the week. Look for: No quantitative results in the abstract and only a 2-to-4-player generalization claim; verify consistency and action-responsiveness beyond cherry-picked rollouts. (Fangfu Liu, Kai He, Tianchang Shen et al., https://arxiv.org/abs/2605.28816)</description>
  </item>
  <item>
    <title>9. Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention</title>
    <link>https://paperfeed.app/2026-W22/arxiv-2605-29548.html</link>
    <guid isPermaLink="false">paperfeed:2026-W22:arxiv:2605.29548</guid>
<pubDate>Sun, 30 Aug 2026 06:15:12 +0000</pubDate>    <description>A concrete, experimentally supported mechanism for why larger models learn rare/complex tasks: reduced gradient interference lets big models allocate enough neurons to frequent tasks that their updates stop overwriting slowly-accumulating rare-task features. Validated with OLMo pretraining from 4M to 4B on controlled tasks. Look for: Evidence is largely synthetic/controlled; be skeptical about how the interference story quantitatively accounts for real-world scaling curves versus other capacity effects. (Jing Huang, Daniel Wurgaft, Rachit Bansal et al., https://arxiv.org/abs/2605.29548)</description>
  </item>
  <item>
    <title>10. When Does LeJEPA Learn a World Model?</title>
    <link>https://paperfeed.app/2026-W22/arxiv-2605-26379.html</link>
    <guid isPermaLink="false">paperfeed:2026-W22:arxiv:2605.26379</guid>
<pubDate>Sun, 30 Aug 2026 06:15:12 +0000</pubDate>    <description>Turns LeJEPA&#39;s empirical recipe into a theorem: alignment-plus-Gaussian-regularization linearly identifies a world&#39;s latent variables under stationary additive-noise dynamics, with Gaussian being the unique compatible latent distribution, and links this to optimal latent-space planning. A rare case of a representation-learning objective getting an identifiability guarantee. Look for: The transition/observation assumptions may be restrictive; judge how much the 1024-dim and pixel-control experiments actually exercise the theory&#39;s boundaries. (David Klindt, Yann LeCun, Randall Balestriero, https://arxiv.org/abs/2605.26379)</description>
  </item>
</channel>
</rss>