Topic: gpu
12 stories found
Today
Perplexity Details Its GPU Embedding Stack: How Ivy, Tulip and ROSE Serve pplx-embed
Perplexity shared details about its GPU-based embedding stack, including Ivy, Tulip, and ROSE, to enhance retrieval quality in AI search products by improving how cheaply embeddings can be run across indexes. This matters because optimizing embedding serving infrastructure can significantly boost the efficiency and performance of AI applications.
Yesterday
Nous Research Adds One-Click Local Model Setup to Hermes Desktop
Nou Research has simplified local model setup in Hermes Desktop to a single click, automatically optimizing and configuring the process for users based on their hardware. This update aims to make AI model deployment more accessible and efficient for end-users.
Tuesday, September 1, 2026
Saturday, August 29, 2026
Friday, August 28, 2026
Vercel AI Open-Sources vgpu: A TypeScript WebGPU Library for AI Agent Shaders
Wednesday, August 26, 2026
Ollama releases: v0.33.1
Ollama released version 0.33.1, which includes updates to Qwen3.8 Flash Next support, cmake patches, and mlxrunner structured output for improved model loading times, highlighting ongoing development and community contributions.
Monday, August 24, 2026

Import AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with Hawkeye
The article discusses three key topics: the ethical considerations regarding machine rights, a new method for automating environment generation using SPADE, and improvements in GPU kernel efficiency through Hawkeye. These advancements highlight the ongoing acceleration in cybersecurity, mathematical computations, and artificial intelligence technologies.
ggml/llama.cpp releases: b10612
The ggml/llama.cpp project released version b10612, which disables a specific test for WebGPU compatibility. This update is important as it enhances the project's support for diverse hardware architectures, including Apple Silicon on macOS.
Sunday, August 23, 2026
ggml/llama.cpp releases: b10594
A code update in ggml/llama.cpp (commit b10594) optimizes device_info handling by skipping unnecessary loops when device information is not printed, reducing overhead, particularly for the CUDA backend. This improvement enhances efficiency without affecting functionality.
12 of 12 items shown. Sources: 107 days indexed.