Paper Feed

Issue 24 · Jun 8–14, 2026

Week 2026-W24

3,588 papers scanned 150 shortlisted 10 picked $10.37 spent

A strong week for efficiency and new framings in transformers (editable/composable KV caches, group-wise sparse attention), plus several assumption-overturning neuroscience results and a genuinely novel BCI platform. Robotics contributes two data-efficiency jumps (human-video-to-dexterous-robot, zero-shot sim-to-real deformables). As always, discount the systems/benchmark overclaims where abstracts report only best-case numbers; the neuroscience picks are preprints, so treat mechanisms as provisional.

  1. AI / ML ✓ read

    Models Take Notes at Prefill: KV Cache Can Be Editable and Composable

    Bojie Li

    Reframes the KV cache as a notebook of memoized, field-conditioned conclusions that can be edited after a correction and RoPE-repositioned/spliced into new contexts, with causal evidence across four model families. If it holds up it is both a new conceptual lens on what prefill computes and a practical serving win (large TTFT reductions, append-only, composes with prefix caching).

    Look for Check the causal claim that the field's own KV drives <1% of the decision, and whether edit+compose stays decision-identical outside the curated benchmarks and for non-CoT settings.

    10 min read · arXiv ↗ ·PDF

  2. BCI ✓ read

    A Fully Endovascular Neural Interface

    Stanton, J., Talei Franzesi, G., Spinazzi, E., Haupt, J. et al.

    A fully endovascular, sub-1-mm3, ultrasound-powered neural implant delivered like a stent, demonstrating autonomic stimulation and blood-pressure modulation in rabbits. This is a genuinely new, less-invasive neural-interface platform rather than an incremental electrode improvement.

    Look for Note that this is stimulation only in a small-animal acute setting; recording capability, chronic safety, and durability remain unproven.

    6 min read · bioRxiv ↗ ·PDF

  3. Robotics ✓ read

    Dexterous Point Policy: Learning Point-based Dexterous Hand Policies from Human Demonstrations

    Beomjun Kim, Seong Hyeon Park, Seunghoon Sim, Seungjun Moon et al.

    Learns dexterous multi-finger manipulation from human videos with zero robot demonstrations by using a shared wrist+fingertip 3D keypoint representation for both observation and action, reporting 75% vs 1% for a VLA baseline. The claim that keypoint-level alignment largely dissolves the human-to-dexterous-robot embodiment gap would be a meaningful data-efficiency jump.

    Look for The task suite and baseline breadth are thin in the abstract; scrutinize how many tasks, how the 1% VLA baseline was configured, and whether keypoints capture contact-rich forces.

    11 min read · arXiv ↗ ·PDF

  4. AI / ML ▲ 2 ✓ read

    The Signs Were Always There: Training-Free Concept Detection and Steering in Raw Transformer Dimensions

    Varun Reddy Nalagatla

    Argues the raw standard basis of transformer hidden states is a training-free, cross-modal feature basis: signs encode content, per-dim reading loses nothing over a full MLP, and flipping sign patterns causally steers concepts. If robust, this undercuts a core premise of SAE/dictionary interpretability and separates reader vs writer dimensions.

    Look for Sweeping claims and only 2 upvotes—verify the steering results and that 'sign alone' truly rivals learned probes rather than reflecting benchmark-specific artifacts.

    11 min read · arXiv ↗ ·PDF

  5. AI / ML ▲ 164 ✓ read

    MiniMax Sparse Attention

    Xunhao Lai, Weiqi Xu, Yufeng Yang, Qiaorui Chen et al.

    A streamlined block-sparse attention over GQA with group-specific top-k selection and a co-designed kernel, reporting 28.4x attention-compute reduction and 14.2x/7.6x prefill/decode speedups at 1M context on a 109B multimodal model. This is the kind of simple, deployable long-context method that tends to get widely adopted.

    Look for Task-level quality parity at 1M tokens is under-reported; check retrieval/agentic quality, not just perplexity, and how it compares to other learned-sparse schemes.

    12 min read · arXiv ↗ ·PDF

  6. Neuroscience ✓ read

    Dynamic trajectory cues drive sequenced integration in approach detectors

    Vashistha, H., Matos, N. C., Wu, H., Clark, D. A.

    Shows luminance change alone evokes an approach/retreat percept in both humans and flies, identifies dual-purpose approach-detector neurons in Drosophila, and finds cues are integrated synergistically only in the natural temporal order. A clean cross-species link from a new percept to a defined, sequence-sensitive circuit computation.

    Look for How strong the causal silencing/imaging evidence is for the 'sequenced integration' claim versus a correlational temporal-order effect.

    9 min read · bioRxiv ↗ ·PDF

  7. Neuroscience ✓ read

    Lineage tracing and live-cell imaging reveal that NeuroD1 does not reprogram microglia into neurons

    Li, X., Li, Y., Cao, Y., Hu, N. et al.

    Using virus-free lineage tracing, longitudinal two-photon imaging, and scRNA-seq, finds NeuroD1 does not convert microglia to neurons and instead drives microglial apoptosis—directly challenging a prominent and contested glia-to-neuron reprogramming literature. Methodologically rigorous negative result that could recontextualize many prior conversion claims.

    Look for The abstract's final sentence appears to contain a contradiction/typo; read the actual lineage-tracing controls to confirm the direction of the claim.

    6 min read · bioRxiv ↗ ·PDF

  8. Neuroscience ✓ read

    A Two-Dimensional Grid-Cell Code for Three-Dimensional Navigation in Freely Flying Bats

    Qi, K. K., Yartsev, M. M.

    Wireless recordings from freely flying bats show grid cells retain a 2D toroidal manifold, and flight paths are organized along transient 2D planes—offering a concrete resolution to how a 2D grid code could support 3D navigation. From the Yartsev lab, this reframes a long-standing debate about grid coding in 3D.

    Look for Whether the 'plane-of-motion' account generalizes beyond structured foraging flights and how robustly the toroidal topology holds during genuinely volumetric maneuvers.

    7 min read · bioRxiv ↗ ·PDF

  9. Robotics ✓ read

    SimWeaver: Zero-Shot RGB Sim-to-Real for Deformable Manipulation

    Wenkang Hu, Haoran Wang, Yitong Li, Liu Liu et al.

    Reports zero-shot RGB sim-to-real for visually complex deformable manipulation (plastic bags, silk) from 200 sim demos per task, with 91% average real success and strong robustness under visual shift where real-data baselines collapse. Deformable RGB sim-to-real without real fine-tuning has been largely unsolved.

    Look for How much comes from the ISP-aware photometric augmentation and measurement-backed simulator; check whether success holds on objects far from the asset-generation distribution.

    8 min read · arXiv ↗ ·PDF

  10. AI / ML ✓ read

    Overcoming State Inertia in Full-Duplex Spoken Language Models via Activation Steering

    Cheng-Kuang Chang, Kai-Wei Chang, Alexander H. Liu, James Glass

    Diagnoses 'state inertia' in full-duplex spoken LMs—internal representations stay biased toward generation just after a barge-in, causing the model to miss the start of user speech—and fixes it training-free via an activation-steering perception vector, with a zero-buffer benchmark and sizable gains. A crisp mechanistic insight into interactive speech models plus a practical intervention.

    Look for Whether the steering vector generalizes across models and conversational conditions, and how latency/quality trade off in real deployment.

    9 min read · arXiv ↗ ·PDF

Also notable

Projects

A strong week for real-time multimodal systems and unconventional systems work: an open streaming video-language interaction stack, agents that write code to reason spatially or compile CUDA kernels, and Apple Neural Engine internals cracked open. Robotics world-model releases were plentiful but mostly early; the standout there is Allen AI's motion-forecasting VLM.

  1. GitHub Speech / Video ★ 1,817 ✓ read

    jd-opensource/JoyAI-VL-Interaction

    JoyAI-VL-Interaction: An Open Real-time Video-Language Interaction System

    An open 8B system that continuously watches video and autonomously decides when to speak, stay silent, or delegate is exactly the proactive real-time multimodal direction the reader tracks. The release is unusually complete — model, training recipe, time-aligned interaction data, quantized checkpoints, and deployment stack — rather than an offline model with a demo video.

    Look for Verify actual end-to-end latency and speak/silence decision quality on your own streams; check whether the interaction data license permits derivative training.

    3 min read ·GitHub ↗ ·Python·Apache-2.0

  2. GitHub AI / ML ★ 371 ✓ read

    NVlabs/SpatialClaw

    SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning

    Replacing rigid tool-calling with a persistent Python kernel the VLM programs against — with segmentation, depth, and geometry tools whose intermediate results it can inspect — is a genuinely fresh action-interface idea. An 11-point average gain across 20 spatial benchmarks and six backbones, training-free, suggests it generalizes rather than overfitting one setup.

    Look for Check inference cost per query (multi-step code execution can be slow/expensive) and whether the gains hold outside the curated benchmark suite.

    4 min read ·GitHub ↗ ·Python

  3. GitHub Tooling ★ 41 ✓ read

    sbryngelson/ANEForge

    Pythonic binding to the Apple Neural Engine

    Pure-ANE execution including on-engine backpropagation and Adam training from ordinary Python, bypassing CoreML entirely, is a capability nobody outside Apple has had. If the MLPerf-valid results hold, it materially changes what Apple-silicon developers can do with the previously opaque Neural Engine.

    Look for It depends on private, unsupported APIs that could break with any macOS update — treat it as research infrastructure, not production, and verify the reported numbers on your own hardware.

    4 min read ·GitHub ↗ ·Python·MIT

  4. GitHub Robotics ★ 134 ✓ read

    allenai/molmo-motion

    Language-conditioned 3D trajectory prediction for arbitrary user-selected points is a new intermediate representation between video models and robot policies, and the demonstrated transfer to both robot planning and motion-guided video generation is compelling. Allen AI releases the model, a million-example corpus, benchmarks, and training recipes.

    Look for Check how well trajectory predictions hold up on cluttered real scenes versus curated evaluation data, and how much the robot-planning transfer depends on downstream machinery.

    3 min read ·GitHub ↗ ·Python·Apache-2.0

  5. GitHub AI / ML ★ 135 ✓ read

    RightNow-AI/AutoMegaKernel

    An agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batch-1 LLM decode, paper: https://arxiv.org/abs/2606.09682

    An agent harness that verifies, fuses, and self-tunes an entire Llama-style decode pass into one persistent CUDA megakernel — retargeting itself across GPU generations — is substantive automated systems engineering, not a wrapper. The honest reporting that its equal-precision bf16 path still loses to cuBLAS makes the int8 wins far more credible.

    Look for Confirm the correctness gating covers your model variant and precision; gains are currently specific to batch-1 int8 decode on inference-class GPUs.

    3 min read ·GitHub ↗ ·Python·MIT

  6. GitHub Speech / Video ★ 70 ✓ read

    HarryHsing/OmniAgent

    OmniAgent (ICML 2026): the first native omni-modal agent for active video perception — a 7B agent that beats Qwen2.5-VL-72B with 73% fewer frames on LVBench.

    A 7B agent that natively decides which frames, audio, or clips to fetch while reasoning — beating a 72B model with 73% fewer frames — is a strong data point that active perception, not brute-force context, is the path for long video understanding. The turn-level RL credit assignment for perception actions is a nice methodological contribution too.

    Look for Traction is minimal and benchmark details are thin; verify the LVBench comparison setup and whether weights and the RL training code are actually released.

    3 min read ·GitHub ↗ ·Python·Apache-2.0

Also notable

The shortlist: top candidates that survived triage · Archive