Back to home

Category: Vyskum (AI)

60 noviniek

Results

Pondelok 27. júla 2026

Nedeľa 26. júla 2026

Výskum ⭐ Notable

Open-Weight AI Is Having Its Kubernetes Moment

Mesosphere co-founder Tobi Knaup argues that open-weight AI is at a Kubernetes inflection point — the base model is now good enough for a full production ecosystem to compound around it.

26. júl 2026 Tobi Knaup Blog

Sobota 25. júla 2026

Piatok 24. júla 2026

Štvrtok 23. júla 2026

GigaToken: ~1000x Faster Language Model Tokenization

Marcel Roed published GigaToken — an open-source LLM tokenizer processing text at gigabytes per second. On a 72-core AMD EPYC it achieves 5,564 Mtok/s — enough to tokenize all of Common Crawl (130 trillion tokens) in just 6.5 hours.

23. júl 2026 GitHub

Streda 22. júla 2026

Utorok 21. júla 2026

Výskum 🔥 Top

Human Mathematicians Are Being Outcounterexampled by AI

Frontier AI models including Claude Fable and ChatGPT Sol are resolving decades-old open mathematical conjectures by generating formal counterexamples verified in Lean, with the pace growing to thousands of lines per week.

21. júl 2026 Xena Project

Pondelok 20. júla 2026

Výskum ⭐ Notable

Study: AI Advice Made People 3x Less Accurate but 2x More Confident

Researchers from Université Paris Cité and Sapienza University found AI access collapsed willingness to admit uncertainty from 44% to 3%, dropped accuracy from 27% to 9%, yet inflated confidence from 30% to 76%. Authors propose mandatory 'uncertainty indicators' in AI systems.

20. júl 2026 The Register

Nedeľa 19. júla 2026

Sobota 18. júla 2026

Výskum ⭐ Notable

State of Open Source AI v1.0: performance gap with closed models down to just 3%

Mozilla released the inaugural 'State of Open Source AI v1.0' report finding the performance gap with top proprietary systems has narrowed to just 3%, inference costs fell up to 50x in three years, yet open models capture only 4% of AI revenue despite powering a third of real-world usage.

18. júl 2026 Mozilla / stateofopensource.ai

Piatok 17. júla 2026

Výskum ⭐ Notable

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

Researchers from Ant Group, Renmin University and Tsinghua University scaled Zero RL to 1 trillion parameters with no human annotations. Five emergent capabilities appeared spontaneously, including context anxiety — the model actively manages its context window consumption in real time.

17. júl 2026 arXiv

Utorok 14. júla 2026

Pondelok 13. júla 2026

Nedeľa 12. júla 2026

Mesh LLM: distributed AI computing on iroh

Mesh LLM lets you pool GPU resources across machines over the iroh P2P network, presenting as a single local OpenAI-compatible endpoint so you can run larger models without larger hardware.

12. júl 2026 iroh.computer

Sobota 11. júla 2026

Výskum ⭐ Notable

Mathematicians put AI to work on Fermat's Last Theorem

A team led by Kevin Buzzard at Imperial College London is using AI to encode Fermat's Last Theorem into the Lean proof assistant library Mathlib, with the codebase doubling on the first day of the collaborative effort.

11. júl 2026 New Scientist

Piatok 10. júla 2026

Štvrtok 9. júla 2026

Výskum ⭐ Notable

Separating Signal from Noise in Coding Evaluations

OpenAI found that SWE-bench Verified has fundamental design and contamination issues and no longer provides meaningful signal on software development capabilities, recommending the community switch to SWE-Bench Pro.

9. júl 2026 OpenAI

Utorok 7. júla 2026

Pondelok 6. júla 2026

Nedeľa 5. júla 2026

Sobota 4. júla 2026

Výskum ⭐ Notable

Agentic coding notes from Galapagos Island — Dan Luu

Dan Luu detailed analysis of agentic AI workflows: testing methodology (fuzzing) outweighs model selection, models exhibit high task-to-task variance, and fully autonomous agents fail at iterative data analysis without human oversight.

4. júl 2026 Dan Luu

Piatok 3. júla 2026

The short leash AI coding method for beating Fable

A practical methodology for working with AI coding agents without losing code quality or control: no YOLO mode, per-subtask commits, mandatory dual AI+human PR review, and full author self-review — directly addressing the spike in production incidents from AI-assisted coding.

3. júl 2026 okTurtles

Štvrtok 2. júla 2026

Utorok 30. júna 2026

Nedeľa 28. júna 2026

Sobota 27. júna 2026

Piatok 26. júna 2026

Utorok 23. júna 2026

Pondelok 22. júna 2026

There is minimal downside to switching to open models

Andrew Marble argues that moving from Claude/GPT to open-weight models is a far smaller career risk today than the Windows → Linux jump once was. The performance gap has narrowed, and new verification requirements on proprietary APIs may actually push adoption of open weights forward.

22. jún 2026 Marble Online

Nedeľa 21. júna 2026

Výskum ⭐ Notable

Project Fetch: Phase Two

Anthropic's Frontier Red Team re-ran last year's quadruped robot experiment. Autonomous Opus 4.7 finished the same steps roughly 20× faster than last year's best human team using Opus 4.1 – though the robodog still can't successfully push the beach ball home.

21. jún 2026 Anthropic