Roomer: Reflective Object-Grounded Model Editing and Repair for 3D Indoor Layout Synthesis Paper • 2608.01973 • Published 2 days ago • 16
UEmbed: Unified Sparse and Dense Multimodal Embeddings Paper • 2608.02583 • Published 2 days ago • 43
Progressive Agent Skill Generation via Reinforcement Learning Paper • 2608.01678 • Published 2 days ago • 52
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks Paper • 2608.01964 • Published 2 days ago • 134
3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering Paper • 2608.01185 • Published 3 days ago • 15
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity Paper • 2608.02603 • Published 2 days ago • 30
In the Driver's Seat: A Multi-Company Study on the Reality of Autonomous Driving System Testing Paper • 2607.15820 • Published 19 days ago • 4
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published 10 days ago • 92
Meshy T2: Fast Native Mesh Generation with Flow Matching Paper • 2607.28675 • Published 8 days ago • 49
Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published 5 days ago • 32
Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability Paper • 2607.26637 • Published 7 days ago • 11
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Paper • 2607.26769 • Published 7 days ago • 24
SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them Paper • 2607.27703 • Published 6 days ago • 22
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Paper • 2607.28509 • Published 6 days ago • 30
Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory Paper • 2607.27919 • Published 6 days ago • 57
PhiZero: A World Model Built Around Physical Language Paper • 2607.28624 • Published 6 days ago • 166