Meta

नवीनतम AI समाचार, मॉडल और रिलीज़ Meta. ['Astryx', 'Business Agent', 'BYT5', 'CodeLlama', 'expressive voice AI', 'FAIRChem v2 UMA', 'Gemma 4', 'Gemma4:e2b', 'HuBERT', 'llama', 'Llama', 'LLaMA', 'LLaMA2-chat-70B', 'Llama 3', 'Llama 3.1', 'Llama 3.1-405B', 'Llama-3.1-405B', 'LLaMA-3.1 405B', 'Llama 3.1-70B', 'Llama 3.1 8B', 'Llama 3.1-8B', 'Llama-3.1-8B', 'Llama 3.2', 'Llama-3.2-1B-Instruct', 'Llama 3.3', 'llama-3.3-70b', 'Llama 3 8B', 'Llama3-8B-Instruct', 'Llama 4', 'Llama 4 Maverick', 'LLaMA 65B', 'Llama 70B', 'M2M100', 'M**a AI', 'mBART', 'Meta AI', 'Meta AI model (model name not specified)', 'Meta Muse Spark', 'Muse Spark', 'Muse Spark 1.1']

शोध 🇺🇸

New DeepSeek-V3 Technical Report: How Hardware-Software Codesign Enables Low-Cost Training of Large Models

A new 14-page technical report from the DeepSeek-V3 team, co-authored by DeepSeek CEO Wenfeng Liang, has been released. The document analyzes the scaling challenges of large language models and the role of hardware infrastructure, proposing solutions based on joint design of models and hardware for cost-effective training and inference.

DeepSeekDeepSeek NVIDIANVIDIA Alibaba/QwenAlibaba/Qwen MetaMeta
Synced27.07 · 18:04
मॉडल 🇺🇸

From GPT-2 to gpt-oss: Analysis of Architectural Improvements and Comparison with Qwen3

OpenAI has released its first open-weight models since 2019 — gpt-oss-120b and gpt-oss-20b. The article examines key architectural changes compared to GPT-2: removal of dropout, replacement of absolute positional embeddings with RoPE, transition from GELU to SwiGLU, introduction of mixture of experts (MoE) and grouped query attention (GQA). It also discusses MXFP4 optimization for running on a single GPU and comparison with Qwen3 and GPT-5.

OpenAIOpenAI Alibaba/QwenAlibaba/Qwen MetaMeta Google/DeepMindGoogle/DeepMind AI21 LabsAI21 Labs TencentTencent
Sebastian Raschka27.07 · 18:03
ताज़ा समाचार