Collections
Discover the best community collections!
Collections including paper arxiv:2504.21853
-
Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play
Paper • 2505.02707 • Published • 85 -
MUSAR: Exploring Multi-Subject Customization from Single-Subject Dataset via Attention Routing
Paper • 2505.02823 • Published • 5 -
PixelHacker: Image Inpainting with Structural and Semantic Consistency
Paper • 2504.20438 • Published • 44 -
Improving Editability in Image Generation with Layer-wise Memory
Paper • 2505.01079 • Published • 29
-
DeepCritic: Deliberate Critique with Large Language Models
Paper • 2505.00662 • Published • 54 -
A Survey of Interactive Generative Video
Paper • 2504.21853 • Published • 46 -
KeySync: A Robust Approach for Leakage-free Lip Synchronization in High Resolution
Paper • 2505.00497 • Published • 17 -
Keysync Demo
📈32Generate synchronized video from audio and video inputs
-
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
Paper • 2504.08641 • Published • 6 -
PRIMA.CPP: Speeding Up 70B-Scale LLM Inference on Low-Resource Everyday Home Clusters
Paper • 2504.08791 • Published • 140 -
Describe Anything: Detailed Localized Image and Video Captioning
Paper • 2504.16072 • Published • 64 -
A Survey of Interactive Generative Video
Paper • 2504.21853 • Published • 46
-
Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
Paper • 2501.18585 • Published • 61 -
RWKV-7 "Goose" with Expressive Dynamic State Evolution
Paper • 2503.14456 • Published • 154 -
DeepMesh: Auto-Regressive Artist-mesh Creation with Reinforcement Learning
Paper • 2503.15265 • Published • 46 -
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
Paper • 2503.15558 • Published • 50
-
Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations
Paper • 2508.09789 • Published • 5 -
MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
Paper • 2508.13186 • Published • 20 -
ZARA: Zero-shot Motion Time-Series Analysis via Knowledge and Retrieval Driven LLM Agents
Paper • 2508.04038 • Published • 1 -
Prompt Orchestration Markup Language
Paper • 2508.13948 • Published • 48
-
Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making Tasks
Paper • 2505.00234 • Published • 26 -
DeepCritic: Deliberate Critique with Large Language Models
Paper • 2505.00662 • Published • 54 -
A Survey of Interactive Generative Video
Paper • 2504.21853 • Published • 46 -
OThink-R1: Intrinsic Fast/Slow Thinking Mode Switching for Over-Reasoning Mitigation
Paper • 2506.02397 • Published • 36
-
Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems
Paper • 2504.01990 • Published • 305 -
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Paper • 2504.10479 • Published • 308 -
What, How, Where, and How Well? A Survey on Test-Time Scaling in Large Language Models
Paper • 2503.24235 • Published • 55 -
Seedream 3.0 Technical Report
Paper • 2504.11346 • Published • 70
-
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
Paper • 2402.04252 • Published • 30 -
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models
Paper • 2402.03749 • Published • 15 -
ScreenAI: A Vision-Language Model for UI and Infographics Understanding
Paper • 2402.04615 • Published • 44 -
EfficientViT-SAM: Accelerated Segment Anything Model Without Performance Loss
Paper • 2402.05008 • Published • 23
-
Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations
Paper • 2508.09789 • Published • 5 -
MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
Paper • 2508.13186 • Published • 20 -
ZARA: Zero-shot Motion Time-Series Analysis via Knowledge and Retrieval Driven LLM Agents
Paper • 2508.04038 • Published • 1 -
Prompt Orchestration Markup Language
Paper • 2508.13948 • Published • 48
-
Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play
Paper • 2505.02707 • Published • 85 -
MUSAR: Exploring Multi-Subject Customization from Single-Subject Dataset via Attention Routing
Paper • 2505.02823 • Published • 5 -
PixelHacker: Image Inpainting with Structural and Semantic Consistency
Paper • 2504.20438 • Published • 44 -
Improving Editability in Image Generation with Layer-wise Memory
Paper • 2505.01079 • Published • 29
-
Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making Tasks
Paper • 2505.00234 • Published • 26 -
DeepCritic: Deliberate Critique with Large Language Models
Paper • 2505.00662 • Published • 54 -
A Survey of Interactive Generative Video
Paper • 2504.21853 • Published • 46 -
OThink-R1: Intrinsic Fast/Slow Thinking Mode Switching for Over-Reasoning Mitigation
Paper • 2506.02397 • Published • 36
-
DeepCritic: Deliberate Critique with Large Language Models
Paper • 2505.00662 • Published • 54 -
A Survey of Interactive Generative Video
Paper • 2504.21853 • Published • 46 -
KeySync: A Robust Approach for Leakage-free Lip Synchronization in High Resolution
Paper • 2505.00497 • Published • 17 -
Keysync Demo
📈32Generate synchronized video from audio and video inputs
-
Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems
Paper • 2504.01990 • Published • 305 -
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Paper • 2504.10479 • Published • 308 -
What, How, Where, and How Well? A Survey on Test-Time Scaling in Large Language Models
Paper • 2503.24235 • Published • 55 -
Seedream 3.0 Technical Report
Paper • 2504.11346 • Published • 70
-
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
Paper • 2504.08641 • Published • 6 -
PRIMA.CPP: Speeding Up 70B-Scale LLM Inference on Low-Resource Everyday Home Clusters
Paper • 2504.08791 • Published • 140 -
Describe Anything: Detailed Localized Image and Video Captioning
Paper • 2504.16072 • Published • 64 -
A Survey of Interactive Generative Video
Paper • 2504.21853 • Published • 46
-
Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
Paper • 2501.18585 • Published • 61 -
RWKV-7 "Goose" with Expressive Dynamic State Evolution
Paper • 2503.14456 • Published • 154 -
DeepMesh: Auto-Regressive Artist-mesh Creation with Reinforcement Learning
Paper • 2503.15265 • Published • 46 -
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
Paper • 2503.15558 • Published • 50
-
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
Paper • 2402.04252 • Published • 30 -
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models
Paper • 2402.03749 • Published • 15 -
ScreenAI: A Vision-Language Model for UI and Infographics Understanding
Paper • 2402.04615 • Published • 44 -
EfficientViT-SAM: Accelerated Segment Anything Model Without Performance Loss
Paper • 2402.05008 • Published • 23