-
7
EMOVA Online Interactive Demo
🔥Live Interactive demo for EMOVA with Qwen-2.5 backbone
-
Emova-ollm/emova-qwen-2-5-3b
Text Generation • Updated • 20 • 2 -
Emova-ollm/emova-qwen-2-5-3b-hf
Feature Extraction • Updated • 31 • 5 -
Emova-ollm/emova-qwen-2-5-7b
Text Generation • Updated • 11 • 1
Collections
Discover the best community collections!
Collections including paper arxiv:2409.18042
-
CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation
Paper • 2410.23090 • Published • 56 -
SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs
Paper • 2410.13276 • Published • 30 -
Personalized Visual Instruction Tuning
Paper • 2410.07113 • Published • 71 -
Differential Transformer
Paper • 2410.05258 • Published • 179
-
impira/layoutlm-document-qa
Document Question Answering • Updated • 16.9k • 1.11k -
EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
Paper • 2409.18042 • Published • 41 -
LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects
Paper • 2504.19838 • Published • 21
-
Building and better understanding vision-language models: insights and future directions
Paper • 2408.12637 • Published • 131 -
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Paper • 2408.11039 • Published • 62 -
Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
Paper • 2408.16725 • Published • 54 -
Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Paper • 2408.15998 • Published • 88
-
RLHF Workflow: From Reward Modeling to Online RLHF
Paper • 2405.07863 • Published • 72 -
Chameleon: Mixed-Modal Early-Fusion Foundation Models
Paper • 2405.09818 • Published • 131 -
Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models
Paper • 2405.15574 • Published • 56 -
An Introduction to Vision-Language Modeling
Paper • 2405.17247 • Published • 90
-
iVideoGPT: Interactive VideoGPTs are Scalable World Models
Paper • 2405.15223 • Published • 17 -
Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models
Paper • 2405.15574 • Published • 56 -
An Introduction to Vision-Language Modeling
Paper • 2405.17247 • Published • 90 -
Matryoshka Multimodal Models
Paper • 2405.17430 • Published • 34