Meta Rolls Out Meta AI in Threads DMs for All Users Globally
Meta rolled out Meta AI to Threads DMs globally on July 27, giving half a billion users private access to the AI chatbot — no more sharing conversations publicly in the feed.
60 noviniek
Meta rolled out Meta AI to Threads DMs globally on July 27, giving half a billion users private access to the AI chatbot — no more sharing conversations publicly in the feed.
Chinese AI lab Moonshot AI released full open weights for Kimi K3, a 2.8-trillion-parameter sparse MoE model. It is the largest open-weight model release in history, natively handling text, image, and video with a 1M token context window.
Nvidia is working on deals totaling more than $750 billion — a $250B backstop for OpenAI's Ohio data center and a $500B+ initiative with SK Group — reviving fears that AI companies are artificially inflating each other's demand.
ChangXin Memory Technologies (CXMT), China's largest DRAM maker, surged 466% on its Shanghai debut after the country's biggest semiconductor IPO ($8.6B), pushing its market cap to $487B and making it mainland China's most valuable listed company.
AMD and Anthropic announced a strategic partnership where AMD commits up to $5 billion in equity investment and will deploy 2 gigawatts of Instinct MI450 GPU compute for Anthropic, with first deployment starting H1 2027.
Google is secretly developing a server AI chip called Frozen v2 that hardwires key parts of Gemini's architecture directly into silicon. Internal sources claim it will deliver 6-10x more AI tokens per watt than current TPUs, targeting deployment in 2028.
SK Group Chairman Chey Tae Won disclosed that Anthropic approached SK Hynix for chip component supplies, signaling Anthropic's push to design and manufacture its own AI semiconductors following OpenAI and Google.
Inflect-Micro-v2 delivers complete text-to-speech synthesis in just 9.36M parameters -- dramatically smaller than standard TTS models -- enabling practical on-device voice generation without cloud dependency.
Anthropic published a detailed guide to context engineering for Claude 5-generation models, covering long contexts, instruction hierarchies, and agent memory management.
Anthropic launched Claude Opus 5, leading the Artificial Analysis leaderboard with major improvements in reasoning, coding, and cybersecurity.
Twenty-five tech companies including Nvidia, Microsoft, and Meta formed a coalition to advocate for development and distribution of open-weight AI models, directly pushing back against closed model approaches.
Moonshot AI will release open weights for the 2.8-trillion-parameter Kimi K3 on July 27, but independent testing revealed a 51% hallucination rate that the company omitted from its benchmark charts.
OpenAI introduced full-duplex voice mode GPT-Live to the ChatGPT desktop app, also enabling voice control of Codex and ChatGPT Work agents during conversations.
Cognition AI, maker of the autonomous coder Devin, acquired startup Poke, expanding its ecosystem of AI tools for developers.
Google expanded access to Gemini Spark, its autonomous AI agent, from the exclusive AI Ultra tier ($100-200/month) to all US AI Pro subscribers ($20/month) on July 24, 2026.
Midjourney, known for AI image generation, acquired popular astrology app Co-Star, moving beyond the visual AI space.
Chinese officials are requesting a more concrete agenda for AI talks with the United States amid ongoing technological and trade tensions.
South Korean President Lee Jae-myung announced a 10-year $880 billion public-private investment plan for semiconductors, AI data centers, and robotics on June 29, 2026, with Samsung and SK Hynix committing to build four new chip fabrication plants.
The White House advisor responsible for energy infrastructure for AI data centers is departing, raising questions about continuity of AI energy policy.
Israel and the United Kingdom appointed dedicated AI ministers, strengthening governmental commitment to strategic AI development and related legislation.
OpenAI launched ChatGPT Work on July 9, 2026 — an agent within ChatGPT that takes over full projects across apps, works autonomously for hours, and returns finished outputs like spreadsheets, slides, documents, and web apps.
AMD unveiled the Instinct MI455X GPU with 432GB HBM4 memory and the Helios rack-scale system delivering 2.9 ExaFLOPS — with OpenAI, Meta, Anthropic, and Microsoft as confirmed customers.
TracerML launched Echo, which dynamically selects and combines open-weight AI models to achieve results comparable to Anthropic Fable at roughly one-third of the cost. It earned 402 upvotes on Hacker News.
Black Forest Labs launched FLUX 3, a multimodal model trained end-to-end across image, video (up to 20 seconds), and audio. Unlike prior pipeline solutions assembling separate models, FLUX 3 is trained jointly across all modalities simultaneously. It launches in limited access.
Etched, an AI chip startup developing specialized ASIC inference chips that operate without traditional GPUs, closed a $300M Series C at a $10.3B valuation. Prior reports indicated negotiations at up to $20B — the actual close confirmed a lower but still substantial figure.
The Cactus team post-trained Google Gemma 4 2B with a 68K-parameter probe that reads hidden states to predict p(wrong), achieving 0.814 AUROC vs. 0.549 for token entropy. Only 15-35% of queries need cloud escalation while matching frontier model performance.
STMicroelectronics lifted its 2026 AI data center revenue forecast to above $1 billion, up from 'nicely above $500 million,' driven by strong demand for 800G and 1.6T pluggable optics and silicon photonics.
Nvidia released a detailed white paper for its Vera data center CPU revealing 88 Olympus cores, 164MB L3 cache, and 1.2 TB/s bandwidth. In SPEC CPU 2026 int rate benchmarks, dual-socket Vera scores 925 vs. 898 for dual-socket AMD EPYC 9755.
TSMC committed $265B to US manufacturing, but US fabs cost 20–50% more than Taiwan, squeezing margins by 2–4%. The costs are expected to be passed on to customers — meaning higher AI chip prices broadly.
Poolside released Laguna S 2.1, a 118B MoE open-weight coding model trained from scratch in 9 weeks, free on Hugging Face — beating DeepSeek-V4-Flash and rivals 10x its size on agentic coding benchmarks.
Nativ is a free, open-source macOS application built natively on Apple MLX for running frontier open-weight models locally on Apple Silicon with no accounts, subscriptions, or cloud dependencies.
Etched, a startup building ASICs specifically optimized for transformer-based AI inference, is reportedly in discussions for a Sequoia-led funding round at valuations potentially reaching $20 billion.
Alibaba unveiled Qwen 3.8-Max-Preview at WAIC 2026 in Shanghai — a 2.4 trillion-parameter multimodal model. The company claims it ranks second only to Claude Fable 5, though no independent benchmarks have been published.
Huawei publicly debuted the Atlas 950 SuperPoD at WAIC 2026 in Shanghai — an 8,192-chip Ascend cluster claiming 6.7× the compute of Nvidia NVL144 and 15× the memory capacity. The full system delivers 1 EFLOPS at FP8 and 256TB of globally addressable memory, with Q4 2026 availability targeted.
OpenAI in Codex 0.144.6 (July 18) corrected the incorrectly advertised 372k tokens to the actual 272k. The developer community simultaneously launched a tracker documenting 35 unscheduled usage-limit resets with no published schedule.
Elon Musk announced xAI is completing training of a 2 trillion parameter model — 33% larger than current Grok 4.5. Release is targeted for August 2026, with speed and token efficiency expected to remain close to its predecessor.
Chinese startup Moonshot AI released Kimi K3, a 2.8-trillion-parameter open-weight model rivaling top US AI offerings, triggering a sharp selloff in semiconductor and AI stocks globally.
Google rolled out connected app integrations for AI Mode Search in the US, letting users build Canva designs, create YouTube Music playlists, or add Instacart grocery items without leaving the search interface.
China's Moonshot AI released Kimi K3, a 2.8 trillion-parameter open-weight model with a 1M context window that outranks Claude Opus 4.8 and ties with GPT-5.6 Sol in Arena benchmarks. Weights release planned for July 27.
Japan's government (METI), Noetra Corp. and NVIDIA launched the world's first national AI infrastructure — an AI factory with 13,750 Vera CPUs and 27,500 Rubin GPUs at 140 MW capacity. Jensen Huang accompanied the announcement and introduced the Cosmos 3 Edge model.
Gemini 3.5 Pro missed its third consecutive launch deadline on July 17, 2026. The fully rebuilt model still fails basic reliability standards. Google is reportedly considering a temporary stopgap release as an interim step.
Thinking Machines Lab, led by former OpenAI CTO Mira Murati and valued at $12B, released its first open-weights model Inkling — a 975B mixture-of-experts model designed for customizability, not raw benchmark performance.
TSMC beat Q2 2026 estimates driven by record AI chip demand. The company raised its full-year revenue growth outlook to over 40% and capital expenditures to $60–64B. It also announced an additional $100B commitment to US manufacturing.
Startup PrismML released Bonsai 27B, the first 27B-parameter model capable of running on iPhone 17 Pro at 11 tokens/second in a 1-bit variant weighing just 3.9 GB. Apple is reportedly evaluating the technology.
Since early June 2026, OpenAI's Codex encrypts instructions passed between the orchestrating agent and its subagents via a new 'multi-agent v2' protocol. Developers building production agent pipelines can no longer audit internal task delegation.
Nvidia has more than halved its approved Asian customer list after introducing a 'white list' requiring tighter compliance checks — including sending staff to customer data centers to verify end users — to prevent AI chips from being diverted to China.
ASML beat Q2 2026 guidance and raised its full-year 2026 revenue outlook for the second time this year to €43–45 billion (from €36–40 billion), driven by AI chip investment and accelerating DRAM orders for AI memory.
TSMC posted record Q2 2026 revenue of $39.6 billion, up 36% year-over-year. AI chips account for roughly a quarter of annual revenue, with 3nm capacity fully booked through Q1 2027.
Analysis revealed Claude Code sends nearly 5x more scaffolding tokens before processing a prompt than OpenCode — a gap that can quadruple with real-world configs.
Ploy.ai documented migrating a production UI-generation agent from Claude Opus 4.8 to GPT-5.6 Sol — 2.2x faster builds and 27% lower cost, but with unexpected behavioral pitfalls.
Comma.ai founder argues AI progress stems primarily from Moore's Law and general compute, not frontier labs overstating their contributions to justify massive fundraising.
Tokyo-based Sakana AI released Fugu, a multi-agent orchestration system packaged as an OpenAI-compatible API that achieves frontier performance by synthesizing outputs from multiple models — without relying on US-controlled models.
Irish data centers now account for 23% of national electricity consumption, driven by AI infrastructure concentration in Dublin, straining the grid and reigniting planning controversies.
Multiple Waymo robotaxis ran out of battery and required towing during July 4 gridlock in San Francisco — city supervisor Bilal Mahmood is demanding a formal inquiry into AV impact on public transit.
ByteDance launched Seedream 5.0 Pro on July 8, a direct rival to OpenAI's GPT-Image 2, featuring native 2K output, editable layer separation, sketch-based editing, and multilingual text rendering in 10+ languages.
IO Fund analyzes how Nvidia finances GPU clouds like CoreWeave and Nebius, which then buy Nvidia chips — a circular structure that inflates revenues on both sides.
Meta launched Muse Spark 1.1 on July 9, 2026, its first commercially paid model priced at $1.25/$4.25 per million tokens. The Superintelligence Labs model offers a 1M context window and agentic capabilities including computer use.
An internal Meta memo reviewed by Reuters revealed plans to begin manufacturing its in-house ASIC chip Iris via TSMC in September 2026. Iris was co-developed with Broadcom and cleared bug-testing in six weeks. Meta aims to double compute from 7 GW to 14 GW and reduce NVIDIA dependence.
OpenAI launched three GPT-5.6 models and ChatGPT Work on July 9, 2026 — an agentic workplace assistant for long-running multi-step tasks, directly competing with Claude Cowork.
Artificial Analysis found Grok 4.5's hallucination rate doubled from 25% to 54% vs. Grok 4.3 even as raw accuracy improved. The model ranked 4th on the Intelligence Index but leads all models on agentic tool-use.