KingNish (Nishith Jain)

posted an update 30 days ago

Post

1065

Wan 2.2 fast upto 10x faster than original wan 2.2

Model: FastVideo/FastWan2.2-TI2V-5B-FullAttn-Diffusers

Space: KingNish/wan2-2-fast

reacted to nicolay-r's post with ❤️ about 1 month ago

Post

2941

📢 For those who planning to start a PhD or research in the UK 🇬🇧 (including AI field in particular) but facing ATAS (Academic Technology Approval Scheme) issues.
Excited to share the ultimate guide for dealing with ATAS refusals and how to write effective rebuttal letters.

🎬 https://youtu.be/bfknM3n-SHs

🔍 From the video you will find:
1. Why appealing an ATAS decision matters even if your visa is approved
2. Which docments to use in understanding the principles behind sponsorship decisions
3. Key tips for proper rebuttal letter structuring

reacted to fdaudens's post with 🚀 about 2 months ago

Post

2233

AudioRAG is becoming real! Just built a demo with ColQwen-Omni that does semantic search on raw audio, no transcription needed.

Drop in a podcast, ask your question, and it finds the exact chunks where it happens. You can also get a written answer.

What’s exciting: it skips transcription, making it faster and better at capturing emotion, ambient sound, and tone, surfacing results text search would miss.

- Demo: fdaudens/colqwen-omni-demo
- Blog post from ColQwen team: https://huggingface.co/blog/manu/colqwen-omni-omnimodal-retrieval

1 reply

·

reacted to Tonic's post with 👍 about 2 months ago

Post

3335

🙋🏻‍♂️ Normalize adding compute & runtime traces to your model cards

2 replies

·

reacted to AdinaY's post with 🔥 about 2 months ago

Post

3424

Kimi-K2 is now available on the hub🔥🚀
This is a trillion-parameter MoE model focused on long context, code, reasoning, and agentic behavior.

moonshotai/kimi-k2-6871243b990f2af5ba60617d

✨ Base & Instruct
✨ 1T total / 32B active - Modified MIT License
✨ 128K context length
✨ Muon optimizer for stable trillion-scale training

1 reply

·

reacted to kanaria007's post with 👀 about 2 months ago

Post

2645

✅ New Article on Hugging Face: Teaching AI to Think Like a System — Not a Toolkit

Title:
🏗️ Understanding Structured Cognitive Architecture: A Unified Framework for AI Reasoning Systems
🔗 Read it here: https://huggingface.co/blog/kanaria007/understanding-structured-cognitive-architecture

Summary:
After exploring how AI can select reasoning modes or learn from failure, this new article zooms out:
*How do all these capabilities form a single mind, not just a menu of functions?*

The **Structured Cognitive Architecture** defines a unified framework where protocols interact coherently — forming a self-organizing, reflective, and ethically grounded reasoning system.

This architecture enables agents to:
• Integrate memory, ethics, reasoning, and identity across layers
• Select and execute reasoning jumps with traceable structure
• Coordinate failure recovery and adaptive learning
• Maintain cross-session identity and self-editing capability

It’s not modular stacking.
It’s **structured systemhood** — cognition with intentional protocol interaction.

Key Features:
• Three-layer design (Foundational, Extended, Learning)
• Semantic rule-layer to avoid protocol interference
• Integrated flow: problem → jump → feedback → pattern update
• Built-in ethics, rollback, and trace integrity

The framework integrates protocols like:
• jump-generator, failure-trace-log, memory-loop, identity-construct
• Extended modules: chronia, structure-cross, evaluation-planning, and more

🧠 Protocol Dataset: kanaria007/agi-structural-intelligence-protocols

Useful for:
• Researchers designing unified AGI architectures
• Developers building reflective protocol-based agents
• Anyone curious how AI can think as a *system*

This isn’t modularity.
It’s **meta-coherence by design**.

2 replies

·

reacted to mlabonne's post with 🔥 about 2 months ago

Post

5462

LiquidAI open-sources a new generation of edge LLMs! 🥳

Based on a new hybrid architecture, these 350M, 700M, and 1.2B models are both fast and performant, ideal for on-device deployment.

I recommend fine-tuning them to power your next edge application. We already provide Colab notebooks to guide you. More to come soon!

📝 Blog post: https://www.liquid.ai/blog/liquid-foundation-models-v2-our-second-series-of-generative-ai-models
🤗 Models: LiquidAI/lfm2-686d721927015b2ad73eaa38

1 reply

·

reacted to a-r-r-o-w's post with 🧠🔥 about 2 months ago

Post

3338

Caching is an essential technique used in diffusion inference serving for speeding up image/video generations. Diffusers just added support for another caching method: First Block Cache - a technique developed by @chengzeyi building upon the ideas of TeaCache.

The idea in short is: if the model predictions do not vary much over successive inference steps, we can skip certain steps where the prediction difference is small. To figure out whether an inference step will make a significant improvement to the overall velocity/noise prediction, we calculate the relative difference of the output of the first transformer block at timestep $t$ with $t-1$, and compare it against a selected threshold. If the difference is lower than the threshold, we skip the step. A higher threshold will lead to more steps being skipped. However, skipping many steps is bad because it can throw off the model predictions, and so we need to test and select the threshold based on level of quality-speed tradeoff for every model we use it with.

Diffusers usage with CogView4:

import torch
from diffusers import CogView4Pipeline
from diffusers.hooks import apply_first_block_cache, FirstBlockCacheConfig

pipe = CogView4Pipeline.from_pretrained("THUDM/CogView4-6B", torch_dtype=torch.bfloat16)
pipe.to("cuda")

apply_first_block_cache(pipe.transformer, FirstBlockCacheConfig(threshold=0.2))

prompt = "A photo of an astronaut riding a horse on mars"
image = pipe(prompt, generator=torch.Generator().manual_seed(42)).images[0]
image.save("output.png")

Below, you'll find the benchmarks and visualizations of the predicted output at different blocks of the Flux DiT.

Docs: https://huggingface.co/docs/diffusers/main/en/optimization/cache
PR: https://github.com/huggingface/diffusers/pull/11180

References:
- First Block Cache: https://github.com/chengzeyi/ParaAttention
- TeaCache: https://github.com/ali-vilab/TeaCache

1 reply

·

reacted to merve's post with 🔥 about 2 months ago

Post

3477

ByteDance released Tar 1.5B and 7B: image-text in image-text out models, fully open-source 👏 ByteDance-Seed/tar-6864cf0d9fe59a3b91cc4260

They have an image tokenizer unified with text, and they de-tokenize using either of two models (LLM and diffusion)
The model is actually a full LLM (Qwen2), the tokenizer converts image tokens 🤯

reacted to AdinaY's post with 🔥 2 months ago

Post

3365

🔥 June highlights from China’s open source ecosystem.

zh-ai-community/june-2025-open-works-from-the-chinese-community-683d66c188f782dc5570ba15

✨Baidu & MiniMax both launched open foundation models
- Baidu: Ernie 4.5 ( from 0.3B -424B ) 🤯
- MiniMax: MiniMax -M1 ( Hybrid MoE reasoning model )

✨Multimodal AI is moving from fusion to full-stack reasoning: unified Any-to-Any pipelines across text, vision, audio, and 3D
- Baidu: ERNIE-4.5-VL-424B
- Moonshot AI: Kimi-VL-A3B
- Alibaba: Ovis-U1
- BAAI: Video-XL-2/OmniGen2
- AntGroup: Ming-Lite-Omni
- Chinese Academy of Science: Stream-Omni
- Bytedance: SeedVR2-3B
- Tencent: Hunyuan 3D 2.1/ SongGeneration
- FishAudio: Openaudio-s1-mini

✨Domain specific models are rapidly emerging
- Alibaba DAMO: Lingshu-7B (medical MLLM)
- BAAI: RoboBrain (Robotics)

✨ So many small models!
- OpenBMB: MiciCPM4 ( on device )
- Qwen: Embedding/Reranker (0.6B)
- Alibaba: Ovis-U1-3B
- Moonshot AI: Kimi-VL-A3B
- Bytedance: SeedVR2-3B

reacted to merve's post with 😎 2 months ago

Post

3056

visual reasoning is now in transformers 🔥
https://huggingface.co/THUDM/GLM-4.1V-9B-Thinking is just released and merged into transformers, we gave it a vibe test run 🤠

it's very good, comes with 64k context length and MIT license 😍
it supports 4k image tokens and any aspect ratio as well!
Notebook: http://colab.research.google.com/drive/1atODIiV57hOZLv16Bjzwd6fwx0yoTorj?usp=sharing
Demo: https://huggingface.co/spaces/THUDM/GLM-4.1V-9B-Thinking-Demo

reacted to Abhaykoul's post with 👀👍🔥 2 months ago

Post

3035

🎉 Dhanishtha 2.0 Preview is Now Open Source!

The world's first Intermediate Thinking Model is now available to everyone!

Dhanishtha 2.0 Preview brings revolutionary intermediate thinking capabilities to the open-source community. Unlike traditional reasoning models that think once, Dhanishtha can think, answer, rethink, answer again, and continue rethinking as needed using multiple blocks between responses.

🚀 Key Features
- Intermediate thinking: Think → Answer → Rethink → Answer → Rethink if needed...
- Token efficient: Uses up to 79% fewer tokens than DeepSeek R1 on similar queries
- Transparent thinking: See the model's reasoning process in real-time
- Open source: Freely available for research and development

HelpingAI/Dhanishtha-2.0-preview
https://helpingai.co/chat

1 reply

·

reacted to eaddario's post with 🚀 2 months ago

Post

3736

Layer-wise and Pruned versions of Qwen/Qwen3-30B-A3B

* Tesor-wise: eaddario/Qwen3-30B-A3B-GGUF
* Pruned: eaddario/Qwen3-30B-A3B-pruned-GGUF

Even though the Perplexity scores of the pruned version are 3 times higher, the ARC, HellaSwag, MMLU, Truthful QA and WinoGrande scores are holding remarkably well, considering two layers were removed (5 and 39). This seems to support Xin Men et al conclusions in
ShortGPT: Layers in Large Language Models are More Redundant Than You Expect (2403.03853)

Results summary in the model's card and test results in the ./scores directory. Questions/feedback is always welcomed.

reacted to Abhaykoul's post with 👍🔥 2 months ago

Post

4511

Introducing Dhanishtha 2.0: World's first Intermediate Thinking Model

Dhanishtha 2.0 is the world's first LLM designed to think between the responses. Unlike other Reasoning LLMs, which think just once.

Dhanishtha can think, rethink, self-evaluate, and refine in between responses using multiple <think> blocks.
This technique makes it Hinghlt Token efficient it Uses up to 79% fewer tokens than DeepSeek R1
---

You can try our model from: https://helpingai.co/chat
Also, we're gonna Open-Source Dhanistha on July 1st.

---
For Devs:
🔑 Get your API key at https://helpingai.co/dashboard

from HelpingAI import HAI  # pip install HelpingAI==1.1.1
from rich import print

hai = HAI(api_key="hl-***********************")

response = hai.chat.completions.create(
    model="Dhanishtha-2.0-preview",
    messages=[{"role": "user", "content": "What is the value of ∫0∞𝑥3/𝑥−1𝑑𝑥 ?"}],
    stream=True,
    hide_think=False # Hide or show models thinking
)

for chunk in response:
    print(chunk.choices[0].delta.content, end="", flush=True)

2 replies

·

reacted to hesamation's post with 🔥 3 months ago

Post

2893

this repo is gold! a collection of LLM apps with multi-agents, MCP, RAG and so much more.

the best way to learn is by building, and this repo provides the blueprint.

Repo: https://github.com/Shubhamsaboo/awesome-llm-apps

posted an update 3 months ago

Post

1104

What's currently the biggest gap in Open Source Datasets ??

4 replies

·

Nishith Jain

AI & ML interests

Recent Activity

Organizations

Nishith Jain

AI & ML interests

Recent Activity

Organizations

KingNish's activity