Abhishek Bhadoriya

abhi22
·

AI & ML interests

- Artificial - Intelligence - Neural Networks - Design - Creativity - Stable diffusion - LLM - RL - Transformer Diffusion - Time Series - Temporal Data - NLP

Recent Activity

View all activity

Organizations

None yet

abhi22's activity

Reacted to singhsidhukuldeep's post with 🔥❤️ 30 days ago
view post
Post
2751
If you have ~300+ GB of V-RAM, you can run Mochi from @genmo

A SOTA model that dramatically closes the gap between closed and open video generation models.

Mochi 1 introduces revolutionary architecture featuring joint reasoning over 44,520 video tokens with full 3D attention. The model implements extended learnable rotary positional embeddings (RoPE) in three dimensions, with network-learned mixing frequencies for space and time axes.

The model incorporates cutting-edge improvements, including:
- SwiGLU feedforward layers
- Query-key normalization for enhanced stability
- Sandwich normalization for controlled internal activations

What is currently available?
The base model delivers impressive 480p video generation with exceptional motion quality and prompt adherence. Released under the Apache 2.0 license, it's freely available for both personal and commercial applications.

What's Coming?
Genmo has announced Mochi 1 HD, scheduled for release later this year, which will feature:
- Enhanced 720p resolution
- Improved motion fidelity
- Better handling of complex scene warping
  • 2 replies
·
updated a collection 9 months ago
updated a collection 9 months ago