Topic: models
70 stories found
Yesterday
Ollama releases: v0.34.0
Ollama released version 0.34.0, allowing users to run their own models in ChatGPT Desktop and improving structured output performance on Apple Silicon, enhancing flexibility and functionality for model users.
Google Launches Agentic Video Understanding for Gemini Flash Models, Cutting Video Tokens by Up to 88%
Google introduced Agentic Video Understanding in Gemini flash models, enabling the platform to navigate videos efficiently rather than processing them at 1 FPS, thereby reducing video tokens by up to 88% and improving performance. This update is significant as it enhances Gemini's efficiency and responsiveness when handling video content.
Friday, September 4, 2026
Where Does Harness-Optimization Value Live? Localized Gains and the Budget-Splitting Trap in Self-Evolving LLM Agents
The article explores how optimizing the "harness" or context around large language models can enhance their performance as autonomous agents. It highlights that while such optimizations can yield localized improvements, they may not always translate to overall budget efficiency, cautioning against over-reliance on budget-splitting strategies for self-evolving LLMs.
Thursday, September 3, 2026
Wednesday, September 2, 2026
Tuesday, September 1, 2026
Monday, August 31, 2026
Friday, August 28, 2026
Show HN: I built a tool showing how AI providers (should) throttle their models
70 of 70 items shown. Sources: 107 days indexed.