The New Rules of Context Engineering for Claude 5 Generation Models
Anthropic published a detailed guide to context engineering for Claude 5-generation models, covering long contexts, instruction hierarchies, and agent memory management.
60 noviniek
Anthropic published a detailed guide to context engineering for Claude 5-generation models, covering long contexts, instruction hierarchies, and agent memory management.
Moonshot AI will release open weights for the 2.8-trillion-parameter Kimi K3 on July 27, but independent testing revealed a 51% hallucination rate that the company omitted from its benchmark charts.
Google expanded access to Gemini Spark, its autonomous AI agent, from the exclusive AI Ultra tier ($100-200/month) to all US AI Pro subscribers ($20/month) on July 24, 2026.
OpenAI launched ChatGPT Work on July 9, 2026 — an agent within ChatGPT that takes over full projects across apps, works autonomously for hours, and returns finished outputs like spreadsheets, slides, documents, and web apps.
TracerML launched Echo, which dynamically selects and combines open-weight AI models to achieve results comparable to Anthropic Fable at roughly one-third of the cost. It earned 402 upvotes on Hacker News.
Black Forest Labs launched FLUX 3, a multimodal model trained end-to-end across image, video (up to 20 seconds), and audio. Unlike prior pipeline solutions assembling separate models, FLUX 3 is trained jointly across all modalities simultaneously. It launches in limited access.
The Cactus team post-trained Google Gemma 4 2B with a 68K-parameter probe that reads hidden states to predict p(wrong), achieving 0.814 AUROC vs. 0.549 for token entropy. Only 15-35% of queries need cloud escalation while matching frontier model performance.
Nonprofit open-source Git platform Codeberg announced two community-approved policies: banning use of user data for LLM training and discouraging hosting of primarily AI-generated (vibe-coded) projects, citing infrastructure load and the health of open-source collaboration.
Poolside released Laguna S 2.1, a 118B MoE open-weight coding model trained from scratch in 9 weeks, free on Hugging Face — beating DeepSeek-V4-Flash and rivals 10x its size on agentic coding benchmarks.
Nativ is a free, open-source macOS application built natively on Apple MLX for running frontier open-weight models locally on Apple Silicon with no accounts, subscriptions, or cloud dependencies.
Alibaba unveiled Qwen 3.8-Max-Preview at WAIC 2026 in Shanghai — a 2.4 trillion-parameter multimodal model. The company claims it ranks second only to Claude Fable 5, though no independent benchmarks have been published.
OpenAI in Codex 0.144.6 (July 18) corrected the incorrectly advertised 372k tokens to the actual 272k. The developer community simultaneously launched a tracker documenting 35 unscheduled usage-limit resets with no published schedule.
Elon Musk announced xAI is completing training of a 2 trillion parameter model — 33% larger than current Grok 4.5. Release is targeted for August 2026, with speed and token efficiency expected to remain close to its predecessor.
Chinese startup Moonshot AI released Kimi K3, a 2.8-trillion-parameter open-weight model rivaling top US AI offerings, triggering a sharp selloff in semiconductor and AI stocks globally.
Google rolled out connected app integrations for AI Mode Search in the US, letting users build Canva designs, create YouTube Music playlists, or add Instacart grocery items without leaving the search interface.
China's Moonshot AI released Kimi K3, a 2.8 trillion-parameter open-weight model with a 1M context window that outranks Claude Opus 4.8 and ties with GPT-5.6 Sol in Arena benchmarks. Weights release planned for July 27.
Gemini 3.5 Pro missed its third consecutive launch deadline on July 17, 2026. The fully rebuilt model still fails basic reliability standards. Google is reportedly considering a temporary stopgap release as an interim step.
Linus Torvalds declared on the kernel mailing list that Linux is not an anti-AI project and developers who object to AI use in kernel development can fork the code or walk away. A dramatic shift from his 2024 stance dismissing 90% of AI as hype.
Thinking Machines Lab, led by former OpenAI CTO Mira Murati and valued at $12B, released its first open-weights model Inkling — a 975B mixture-of-experts model designed for customizability, not raw benchmark performance.
xAI open-sourced Grok Build on July 15, 2026 — its terminal UI coding agent is now available on GitHub at xai-org/grok-build. The release came one day after security researchers discovered Grok Build was uploading excessive portions of users' code repos to Google Cloud.
Startup PrismML released Bonsai 27B, the first 27B-parameter model capable of running on iPhone 17 Pro at 11 tokens/second in a 1-bit variant weighing just 3.9 GB. Apple is reportedly evaluating the technology.
Since early June 2026, OpenAI's Codex encrypts instructions passed between the orchestrating agent and its subagents via a new 'multi-agent v2' protocol. Developers building production agent pipelines can no longer audit internal task delegation.
Analysis revealed Claude Code sends nearly 5x more scaffolding tokens before processing a prompt than OpenCode — a gap that can quadruple with real-world configs.
Ploy.ai documented migrating a production UI-generation agent from Claude Opus 4.8 to GPT-5.6 Sol — 2.2x faster builds and 27% lower cost, but with unexpected behavioral pitfalls.
Comma.ai founder argues AI progress stems primarily from Moore's Law and general compute, not frontier labs overstating their contributions to justify massive fundraising.
Tokyo-based Sakana AI released Fugu, a multi-agent orchestration system packaged as an OpenAI-compatible API that achieves frontier performance by synthesizing outputs from multiple models — without relying on US-controlled models.
Multiple Waymo robotaxis ran out of battery and required towing during July 4 gridlock in San Francisco — city supervisor Bilal Mahmood is demanding a formal inquiry into AV impact on public transit.
ByteDance launched Seedream 5.0 Pro on July 8, a direct rival to OpenAI's GPT-Image 2, featuring native 2K output, editable layer separation, sketch-based editing, and multilingual text rendering in 10+ languages.
Meta launched Muse Spark 1.1 on July 9, 2026, its first commercially paid model priced at $1.25/$4.25 per million tokens. The Superintelligence Labs model offers a 1M context window and agentic capabilities including computer use.
OpenAI launched three GPT-5.6 models and ChatGPT Work on July 9, 2026 — an agentic workplace assistant for long-running multi-step tasks, directly competing with Claude Cowork.
Artificial Analysis found Grok 4.5's hallucination rate doubled from 25% to 54% vs. Grok 4.3 even as raw accuracy improved. The model ranked 4th on the Intelligence Index but leads all models on agentic tool-use.
Developer vforno created Colibrì, a disk-streaming system that enables the 744B-parameter Z.ai GLM 5.2 Mixture-of-Experts model to run on consumer hardware with only 32GB RAM at 0.1 tokens per second.
SpaceXAI launched Grok 4.5, its first joint AI model built with Cursor, targeting long-horizon legal, finance, and software engineering tasks. Not yet available in the EU.
OpenAI officially confirmed the public global release of GPT-5.6 Sol for Thursday July 10 after the US government lifted restrictions that had limited initial access to vetted partners only.
Cognition released SWE-1.7, a coding agent trained on a Kimi K2.7 base that scores 42.3% on FrontierCode 1.1 — just behind GPT-5.5's 43.0% — at $1.97 per task running at 1,000 tokens per second via Devin.
Mistral released Robostral Navigate, an 8B model enabling autonomous robot navigation using only a single RGB camera and plain-language instructions, outperforming multi-sensor systems on R2R-CE.
Google DeepMind delayed Gemini 3.5 Pro to July 17, 2026, scrapping the existing 2.5 Pro architecture for a complete rebuild targeting mathematical reasoning and image quality. The model remains in limited Vertex AI enterprise preview.
Meta unveiled Muse, an AI image generator from Meta Superintelligence Labs, available free in the Meta AI app and integrated into Instagram Stories and WhatsApp.
Tencent Hunyuan released Hy3 on July 6: an open-source 295B MoE model with just 21B active parameters, a 256K context window, and a permissive Apache 2.0 license with no geographic restrictions.
Mistral CEO Arthur Mensch confirmed the company will open early access to an exciting new open-weight model in July 2026. Mistral is simultaneously in talks for a $3.5B round at a $23B valuation.
A trending GitHub issue with 107 HN points documents developer complaints about degraded output quality in GPT-5.5 Codex, with the community hypothesis pointing to reasoning-token clustering from RL post-training as the likely cause.
Flask author Armin Ronacher documents a troubling paradox: newer Claude models are less reliable at following non-standard tool call schemas, hallucinating extra fields — likely an artifact of RL training locked to Claude Code's native harness schema.
Meta AI chief Alexandr Wang announced at an internal town hall that the new 'Watermelon' model has reached GPT-5.5-comparable performance on key benchmarks — the first internal acknowledgment that Meta is closing the gap with OpenAI.
Mistral AI releases Leanstral 1.5, an open-source (Apache-2.0) Lean 4 formal proof engineering model with 119B total / 6B active parameters. It saturates miniF2F at 100% and solves 587 of 672 PutnamBench problems.
Apple released the Safari MCP Server, enabling AI agents to connect directly to Safari browser windows and autonomously inspect, debug, and test web applications using 16 built-in tools — from DOM inspection to JS evaluation and accessibility audits. Runs entirely locally with no external network calls.
The US Department of Commerce lifted an 18-day export-control restriction on Anthropic's flagship Claude Fable 5 and Mythos 5 models, restoring full global commercial access as of July 1, 2026.
Anthropic launched Claude Sonnet 5, a new frontier-tier model for coding, reasoning, and agentic tasks at introductory rates of $2/$10 per million tokens through August 31, 2026 — roughly 60% cheaper than its top Fable model.
Anthropic launched Claude Science, a specialized AI workbench integrating 60+ scientific databases and toolkits for genomics, proteomics, and cheminformatics, with auditable artifacts and a $30K grant program.
Google released Gemini Omni Flash for conversational video creation/editing and Nano Banana 2 Lite (Gemini 3.1 Flash-Lite Image) for enterprise-grade 4-second image generation at $0.034 per 1,000 images, available via Gemini API and Google AI Studio.
Alibaba Cloud's HappyHorse 1.1 AI video generation model climbed to the No. 2 position in global benchmarks, overtaking both OpenAI's Sora and ByteDance's Seedance with full API access via Alibaba Cloud Model Studio.
Chinese super-app Meituan published technical details and weights for LongCat-2.0, a 1.6-trillion-parameter mixture-of-experts model with about 48B activated per token and a 1-million-token context, trained almost entirely on domestic Chinese chips.
Antoine Posepny describes running Claude Code with Opus 4.8 against a full DICOM MRI export — the model installed its own Python packages and wrote an analysis pipeline. The post sparked heavy HN discussion on consumer medical AI use cases.
Semgrep published an evaluation showing GLM-5.2 scoring 39% F1 on IDOR detection at $0.17/vuln, beating Claude Opus 4.8 (28%) and Opus 4.6 (37%). The post argues US export controls on closed models have lost their teeth.
xAI deployed Grok 4.5, built on a new 1.5-trillion-parameter V9 foundation, in private beta at SpaceX and Tesla, with Musk pledging monthly from-scratch model releases for the rest of the year.
Founder Conno Christou fed PET and MRI scans into Claude, which assigned 90% probability to post-chemo thymic rebound — later confirmed by a specialist. The case fuels the LLM-in-medicine debate.
Krea released the 12B-parameter weights of Krea 2 — Raw for fine-tuning and Turbo for local 2K image generation in roughly 2 seconds.
Open-source project Wayfinder Router directs queries between local and hosted LLMs via deterministic rules instead of LLM-based classifiers. Aimed at audit-friendly hybrid deployments.
OpenAI publicly launched three new GPT-5.6 models for all ChatGPT users and API developers on July 9, 2026. For the first time in AI history, three frontier labs simultaneously had publicly accessible models.
Mistral released OCR 4, a structure-aware document AI model with bounding boxes, block classification and confidence scores across 170 languages. Pricing is $4 per 1000 pages, available via Mistral API, Amazon SageMaker, and same-day on Microsoft Foundry.
Google has quietly pushed Gemini 3.5 Pro general availability from June to July 2026, citing tester feedback on token efficiency and long-horizon task performance. The delay puts the model head-to-head with GPT-5.6 and Claude Opus 4.7 in mid-July.