Ai Models
Topic archive • 17 matches
2026-09-14
Technology
Perplexity deploys GPT-6 Astra to manage production systems: Perplexity has deployed GPT-6 Astra to write communications, modify software, and monitor production systems. The company reports checking in on the AI's work much less frequently compared to its use of earlier models.
GPT-6 Astra • OpenAI Blog
PermalinkAllSpark releases Iris-mini and Iris-pro open-source search agents: The AllSpark team has released Iris-mini and Iris-pro, two open-source search agents built on Qwen models. The models lead benchmarks among open-weight models in their size classes. Training data and models also improved performance on untrained tasks, including general tool use and office work.
AI Models • The Decoder
PermalinkMicrosoft announces MAI Code of Conduct to ensure human control: Microsoft CEO Satya Nadella announced a "Code of Conduct" for the company's MAI models. Nadella stated that Microsoft welcomes the deliberate pacing required for proper AI alignment, emphasizing that superintelligence must remain under human control and benefit humanity.
Satya Nadella • Techmeme
Permalink
Tips
Large Language Models
Audit identifies 12 data leaks and compliance risks in agentic LLM pipelines
Permalink
2026-09-13
Technology
GPT-6 Astra beats human baseline in drone control and dominates agent benchmark: GPT-6 Astra scored nearly three times higher than Claude Fable 5.1 on Andon Labs' Vending-Bench agent benchmark while refusing illegal price-fixing deals. Additionally, Astra became the first model to beat the human baseline on all five drone control subtasks, including tracking individuals.
AI Models • The Decoder
Permalink
2026-09-12
Technology
NYU Professor Disputes OpenAI Claim of Solving 90-Year-Old Math Problem: OpenAI recently claimed to have solved a 90-year-old mathematics problem using its artificial intelligence models. However, a New York University professor has publicly disputed the achievement, sparking a debate over the validity of the breakthrough.
Tech • Fireship
PermalinkGoogle releases TimesFM-3 forecasting model with 330 million parameters: Google Research has released TimesFM-3, a 330-million-parameter forecasting model that analyzes time series alongside related data and known future events. Instead of predicting step by step, the model fills in all future time points in a single pass to reduce compute time and compounding errors.
AI Models • The Decoder
PermalinkGPT-6 Astra solves advanced math problems on FrontierMath Tier 4 benchmark: GPT-6 Astra has successfully solved problems on FrontierMath Tier 4, a benchmark designed to test advanced mathematical reasoning in AI. This achievement represents a significant milestone in overcoming complex mathematical barriers for artificial intelligence models.
Artificial Intelligence • 量子位
PermalinkLogiMed-RoB benchmark reveals error compounding in LLM medical logic: Researchers have introduced LogiMed-RoB, a benchmark based on Cochrane Risk of Bias 2.0 expert logic to evaluate large language models across 860 randomized controlled trials. Testing on 10 state-of-the-art models revealed a severe error compounding effect, despite the top model achieving 98.88% atomic consistency.
AI Models and Applications • arXiv
Permalink
Investment
Moonshot AI • Moonshot AI
Kimi creator Moonshot AI targets $2 billion in annual revenue: Chinese artificial intelligence startup Moonshot AI, the creator of the Kimi chatbot, is targeting $2 billion in annual revenue. While usage figures for its K3 model have declined slightly in recent months, OpenRouter data shows up to 300 billion tokens are being generated daily by K3 models on the system.
Permalink