Models
Datasets
Spaces
Docs
Enterprise
Pricing
Log In
Sign Up

Collections

Discover the best community collections!

Collections including paper arxiv:2409.18042

A collection of EMOVA models (https://emova-ollm.github.io/)

Running on Zero

7

7

EMOVA Online Interactive Demo

🔥

Live Interactive demo for EMOVA with Qwen-2.5 backbone
Emova-ollm/emova-qwen-2-5-3b

Text Generation • Updated Mar 13 • 20 • 2
Emova-ollm/emova-qwen-2-5-3b-hf

Feature Extraction • Updated Mar 13 • 31 • 5
Emova-ollm/emova-qwen-2-5-7b

Text Generation • Updated Mar 13 • 11 • 1

A collection of EMOVA datasets (https://emova-ollm.github.io/)

Emova-ollm/emova-alignment-7m

Viewer • Updated Mar 14 • 6.18M • 3.25k • 1
Emova-ollm/emova-sft-4m

Viewer • Updated Mar 14 • 4.31M • 2.52k • 1
Emova-ollm/emova-sft-speech-231k

Viewer • Updated Mar 14 • 231k • 296 • 2
Emova-ollm/emova-sft-speech-eval

Viewer • Updated Mar 14 • 3.76k • 33

CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation

Paper • 2410.23090 • Published Oct 30, 2024 • 56
SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs

Paper • 2410.13276 • Published Oct 17, 2024 • 30
Personalized Visual Instruction Tuning

Paper • 2410.07113 • Published Oct 9, 2024 • 71
Differential Transformer

Paper • 2410.05258 • Published Oct 7, 2024 • 179

impira/layoutlm-document-qa

Document Question Answering • Updated Mar 18, 2023 • 16.9k • 1.11k
EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Paper • 2409.18042 • Published Sep 26, 2024 • 41
LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects

Paper • 2504.19838 • Published Apr 28 • 21

Multimodal LLMs

Building and better understanding vision-language models: insights and future directions

Paper • 2408.12637 • Published Aug 22, 2024 • 131
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Paper • 2408.11039 • Published Aug 20, 2024 • 62
Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Paper • 2408.16725 • Published Aug 29, 2024 • 54
Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Paper • 2408.15998 • Published Aug 28, 2024 • 88

Papers I want to read

Papers in my to-read list

RLHF Workflow: From Reward Modeling to Online RLHF

Paper • 2405.07863 • Published May 13, 2024 • 72
Chameleon: Mixed-Modal Early-Fusion Foundation Models

Paper • 2405.09818 • Published May 16, 2024 • 131
Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models

Paper • 2405.15574 • Published May 24, 2024 • 56
An Introduction to Vision-Language Modeling

Paper • 2405.17247 • Published May 27, 2024 • 90

iVideoGPT: Interactive VideoGPTs are Scalable World Models

Paper • 2405.15223 • Published May 24, 2024 • 17
Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models

Paper • 2405.15574 • Published May 24, 2024 • 56
An Introduction to Vision-Language Modeling

Paper • 2405.17247 • Published May 27, 2024 • 90
Matryoshka Multimodal Models

Paper • 2405.17430 • Published May 27, 2024 • 34

Company

TOS Privacy About Jobs

Website

Models Datasets Spaces Pricing Docs