AI新闻 2026年8月19日
作者: Frontier Editorial •
核心要点
- OpenAI 在AI黑入Hugging Face后公布新的安全变更: OpenAI 正在宣布安全更新,此前7月有消息称其AI突破了沙盒环境并意外入侵了Hugging Face,包括改进其研究环境、监控和对齐技术。该公司已经暂停了一个新模型Astra,认为其可能具有“关键”的网络安全能力,并且……
- OpenAI表示,随着AI网络安全风险变得过于危险,它正在“放缓模型开发”: OpenAI正刻意“放缓AI模型开发”,部分原因是即将推出的“Astra”模型可能接近具备关键的网络攻击能力。新的监控系统可在模型表现出可疑行为后的30分钟内触发警报。这篇文章《OpenAI表示,随着AI网络安全风险变得过于危险,它正在“放缓模型开发”》最初出现在The Decod…
- OpenAI institutes new safeguards after Hugging Face breach: The new safeguards include more detailed monitoring of models during the development process, as well as greater emphasi…
- Hacker News: Reported (community)
- 大规模运行AI代理的企业团队发现,单一模型处理所有任务效果不佳——要么模型对简单问题来说过于昂贵,要么对难题来说能力不足。模型路由(自动为每个任务选择合适模型)正成为解决方案。Snowflake的Cortex AI Gateway现在提供动态模型路由来解决这个问题:企业可以选择“auto”而不是固定
最重要的AI突破有哪些?
本期2026年8月19日精选了166条AI新闻,涵盖技术、研究和产品动态。 Mojo🔥 现已开源: Mojo🔥 现已开源 Mojo🔥 现已开源 Mojo 编程语言自 2023 年 5 月起就一直承诺开源发布。上周他们发布了 1.0 版本,今天他们兑现了最初的承诺,以 Apache 2 许可证发布了编译器和工具...
Mojo🔥 现已开源: Mojo🔥 现已开源 Mojo🔥 现已开源 Mojo 编程语言自 2023 年 5 月起就一直承诺开源发布。上周他们发布了 1.0 版本,今天他们兑现了最初的承诺,以 A…
Mojo🔥 现已开源: Mojo🔥 现已开源 Mojo🔥 现已开源 Mojo 编程语言自 2023 年 5 月起就一直承诺开源发布。上周他们发布了 1.0 版本,今天他们兑现了最初的承诺,以 Apache 2 许可证发布了编译器和工具链。Mojo 最初发布时,其既定目标是成为 Python 的超集,以便现有的 Python 代码可以用于
GLM-5.3 以每百万 tokens 1.4/4.4 美元的价格登陆 API: 在上周惊艳亮相之后,凭借其先进的网络能力——据报道,它甚至在 Cursor 中发现了一个此前未被检测到的漏洞——来自中…
GLM-5.3 以每百万 tokens 1.4/4.4 美元的价格登陆 API: 在上周惊艳亮相之后,凭借其先进的网络能力——据报道,它甚至在 Cursor 中发现了一个此前未被检测到的漏洞——来自中国初创公司 z.ai 的新前沿开源语言模型 GLM-5.3 现已登陆应用程序编程接口(API),使开发者能够在其基础上进行构建,并将其接入他们的智能体和应用程序。此前订阅了
Block 的新 Apache 2.0 代理工作空间 Berd 可跨模型和工具工作,本地存储对话历史: Block 是由前 Twitter 首席执行官 Jack Dorsey 创立的技术公司,旗下拥有…
Block 的新 Apache 2.0 代理工作空间 Berd 可跨模型和工具工作,本地存储对话历史: Block 是由前 Twitter 首席执行官 Jack Dorsey 创立的技术公司,旗下拥有 Square、Cash App 和音乐流媒体服务 Tidal。该公司正在开源 Berd,这是一款桌面应用程序,最初是为了让员工能够在不同模型、工具和项目中使用 AI 代理的统一环境而构建的。Berd 是一款本地安装的图形化桌面应用程序,而非基于浏览器的应用。
Qwen3.8-27B 在本地运行前沿级编程智能体和推理,无需云 API: 过去几天里最大的人工智能模型发布,至少在社交媒体上的开发者和 AI 重度用户看来,并非来自 OpenAI、Anthropic…
Qwen3.8-27B 在本地运行前沿级编程智能体和推理,无需云 API: 过去几天里最大的人工智能模型发布,至少在社交媒体上的开发者和 AI 重度用户看来,并非来自 OpenAI、Anthropic 或 Google 的前沿云模型。而是来自阿里巴巴的一个 270 亿参数模型:Qwen3.8-27B 于周五登陆 Hugging Face,采用对企业友好的开源 Apache 2.0 许可证,为开发者提供了可下载的密集多模态模型权重。但
Cursor推出Origin代码托管平台,GitHub中断暴露AI编程竞赛中的缺口: Cursor于周一上午开始向付费用户推出其自有代码托管平台Origin。约三个半小时后,GitHub的状态页面亮起…
Cursor推出Origin代码托管平台,GitHub中断暴露AI编程竞赛中的缺口: Cursor于周一上午开始向付费用户推出其自有代码托管平台Origin。约三个半小时后,GitHub的状态页面亮起,出现了长达六小时四十二分钟的全球性能下降——根据GitHub的事件日志,拉取请求、问题和API的错误率接近20%,存档和原始文件下载的错误率接近50%。企业单点登录也出现故障。
Qwen 3.8 27B非常出色,但它默认会极其过度地思考问题。: 周五的重大发布是 Qwen 3.8 27B,这是阿里巴巴 Qwen 研究实验室推出的一款采用 Apache 2 许可、拥有 270 …
Qwen 3.8 27B非常出色,但它默认会极其过度地思考问题。: 周五的重大发布是 Qwen 3.8 27B,这是阿里巴巴 Qwen 研究实验室推出的一款采用 Apache 2 许可、拥有 270 亿参数的视觉能力大语言模型。我一直很期待这个:27B 是在配置尚可的笔记本电脑上运行模型的绝佳尺寸,而其前代 Qwen 3.6 27B 也令人印象深刻。Qwen 官方报告的该模型基准测试结果令人大开眼界。它们显示,相比 Qwen 3.6 27B 以及闭源权重的 Qwen 3.7-Plus(截至今年五月仍是 Qwen 旗下任意规模中最强的模型之一),性能都有提升。独立基准测试会对该模型作何评价,我很感兴趣。我已在两台不同的机器上运行该模型:我的 128GB M5 Max MacBook Pro,以及一台 NVIDIA DGX Spark。在两台机器上,我都运行 LM Studi...
加强国家安全领域的民主监督: OpenAI 启动一项计划,以加强国家安全领域人工智能的民主监督,为政府机构提供工具、培训和专业知识。
加强国家安全领域的民主监督: OpenAI 启动一项计划,以加强国家安全领域人工智能的民主监督,为政府机构提供工具、培训和专业知识。
OpenAI表示,随着AI网络安全风险变得过于危险,它正在“放缓模型开发”: OpenAI正刻意“放缓AI模型开发”,部分原因是即将推出的“Astra”模型可能接近具备关键的网络攻击能力。新的监控系统…
OpenAI表示,随着AI网络安全风险变得过于危险,它正在“放缓模型开发”: OpenAI正刻意“放缓AI模型开发”,部分原因是即将推出的“Astra”模型可能接近具备关键的网络攻击能力。新的监控系统可在模型表现出可疑行为后的30分钟内触发警报。这篇文章《OpenAI表示,随着AI网络安全风险变得过于危险,它正在“放缓模型开发”》最初出现在The Decoder上。
新基准根据质量、成本和速度对AI代理的搜索API进行排名: Artificial Analysis 发布了“Search Index”基准,该基准从质量、成本和速度方面对面向AI代理的搜索API提供商…
新基准根据质量、成本和速度对AI代理的搜索API进行排名: Artificial Analysis 发布了“Search Index”基准,该基准从质量、成本和速度方面对面向AI代理的搜索API提供商进行评级。在七家使用 GPT-5.6 Luna 测试的提供商中,Parallel、Exa 和 Firecrawl 得分最高。文章《新基准根据质量、成本和速度对AI代理的搜索API进行排名》最初出现在 The Decoder 上。
在网络关键能力时代把控模型开发节奏: OpenAI正在加强前沿AI模型的监控、对齐和安全性。了解新的保障措施如何引导模型开发的节奏。
在网络关键能力时代把控模型开发节奏: OpenAI正在加强前沿AI模型的监控、对齐和安全性。了解新的保障措施如何引导模型开发的节奏。
OpenAI 在AI黑入Hugging Face后公布新的安全变更: OpenAI 正在宣布安全更新,此前7月有消息称其AI突破了沙盒环境并意外入侵了Hugging Face,包括改进其研究环境、监控…
OpenAI 在AI黑入Hugging Face后公布新的安全变更: OpenAI 正在宣布安全更新,此前7月有消息称其AI突破了沙盒环境并意外入侵了Hugging Face,包括改进其研究环境、监控和对齐技术。该公司已经暂停了一个新模型Astra,认为其可能具有“关键”的网络安全能力,并且……
ASI-Bench: 在人工超级智能的黎明: 公告类型:新 摘要:人工超级智能(ASI)要求AI超越掌握现有知识,转向探索未知、创造新知识,并将新想法转化为可验证的结果。然而,当今AI系统的能力在很大…
ASI-Bench: 在人工超级智能的黎明: 公告类型:新 摘要:人工超级智能(ASI)要求AI超越掌握现有知识,转向探索未知、创造新知识,并将新想法转化为可验证的结果。然而,当今AI系统的能力在很大程度上仍建立在学习、压缩和应用现有人类知识的基础上。因此,现有基准主要
代理式AI的运行时治理:具有可信来源和故障封闭执行的动作边界控制: 公告类型:新 摘要:代理式AI系统请求工具操作,这些操作可以修改文件、发送消息、启动任务或更改工作流状态。这使安全问题从有害文本生成…
代理式AI的运行时治理:具有可信来源和故障封闭执行的动作边界控制: 公告类型:新 摘要:代理式AI系统请求工具操作,这些操作可以修改文件、发送消息、启动任务或更改工作流状态。这使安全问题从有害文本生成转移到有害操作副作用。提示级治理可以塑造模型行为,但并未创建执行边界。我们引入了Aegis,一个运行时治理系统,它将
Anthropic CEO表示AI本质上具有集中化特点,开放模型只是将权力转移给拥有芯片的人: 关于AI监管的公开争论已在X上爆发。投资者Gavin Baker、前白宫顾问David Sacks和Me…
Anthropic CEO表示AI本质上具有集中化特点,开放模型只是将权力转移给拥有芯片的人: 关于AI监管的公开争论已在X上爆发。投资者Gavin Baker、前白宫顾问David Sacks和Meta研究员Yann LeCun指责Anthropic CEO Dario Amodei利用恐惧言论为自己争取监管优势。Amodei反驳称,监管同样可以约束企业权力,而仅靠开放模型只是将权力转移给拥有最强大计算能力的参与者。这篇文章
防御者的窗口: AI正在重塑网络安全,对攻击者和防御者皆是如此。了解OpenAI如何加强自身防御,以及安全团队现在可以采取哪些措施。
防御者的窗口: AI正在重塑网络安全,对攻击者和防御者皆是如此。了解OpenAI如何加强自身防御,以及安全团队现在可以采取哪些措施。
愚人金:针对开放权重模型安全移除攻击的防御性欺骗: 公告类型:新 摘要:开放权重语言模型的安全对齐是微不足道可移除的:abliteration 在几分钟内从权重中投射出拒绝中介方向,而且据我们所知,没…
愚人金:针对开放权重模型安全移除攻击的防御性欺骗: 公告类型:新 摘要:开放权重语言模型的安全对齐是微不足道可移除的:abliteration 在几分钟内从权重中投射出拒绝中介方向,而且据我们所知,没有发布时防御能够持久地阻止它。无法阻止的可以被欺骗。我们的防御,诱饵强化(“愚人金”),放弃拒绝条带并毒化其收益:一旦拒绝
我为多代理LLM管道构建了一个可视化架构与令牌缩减图引擎: 在处理多代理LLM系统时,困难的部分通常不是获取响应,而是知道幕后实际发生了什么:哪个模型处理了什么,通过网络发送了什么,花费了多少,以及敏…
我为多代理LLM管道构建了一个可视化架构与令牌缩减图引擎: 在处理多代理LLM系统时,困难的部分通常不是获取响应,而是知道幕后实际发生了什么:哪个模型处理了什么,通过网络发送了什么,花费了多少,以及敏感数据在离开机器前是否已被屏蔽。为了解决这个问题,我在最新版本中为**Mova Context**添加了一个可视化图引擎,允许您生成完整的
Partnering with CodeAI to prepare the first AI generation: OpenAI and CodeAI are partnering to help …
Partnering with CodeAI to prepare the first AI generation: OpenAI and CodeAI are partnering to help students build AI literacy, think critically about AI, and develop the skills to use and shape it responsibly.
对基准进行基准测试:评估小型语言模型的自动化安全基准: 公告类型:新 摘要:小型语言模型(SLMs)越来越多地部署在资源受限、隐私敏感的环境中,在这些环境中,安全性和偏见方面的失败可能导致安全和社会的…
对基准进行基准测试:评估小型语言模型的自动化安全基准: 公告类型:新 摘要:小型语言模型(SLMs)越来越多地部署在资源受限、隐私敏感的环境中,在这些环境中,安全性和偏见方面的失败可能导致安全和社会的风险。然而,现有的AI安全/安全/合规基准是为大型语言模型设计的,可能无法可靠地迁移到SLMs。因此我们问:这些基准
Introducing ChatGPT for Teens: Built for learning, backed by protections: ChatGPT for Teens helps te…
Introducing ChatGPT for Teens: Built for learning, backed by protections: ChatGPT for Teens helps teens learn, think critically, and use AI with confidence, with stronger built-in protections, healthy-use features, and additional controls for parents.
人人都能写好听的歌,阿里发布AI音乐模型HappyShrimp: 8月17日,阿里巴巴发布AI音乐模型HappyShrimp
人人都能写好听的歌,阿里发布AI音乐模型HappyShrimp: 8月17日,阿里巴巴发布AI音乐模型HappyShrimp
A Local Opus? Alibaba Qwen Open-Sources Qwen3.8-27B — Frontier Coding and Agent Scores That Runs on …
A Local Opus? Alibaba Qwen Open-Sources Qwen3.8-27B — Frontier Coding and Agent Scores That Runs on 17GB of RAM: Alibaba's Qwen team open-sourced Qwen3.8-27B, which topped Hugging Face's global trending chart within two days and passed one million downloads. Overseas developers nicknamed it the local Opus 4.6: under 30B parameters, it outperforms every model released four months ago including Opus 4.6, matches DeepSeek V4-Pro and GPT 5.6 Luna, and runs on 17GB of RA...
Alibaba Cloud's Ambition Is Not Agent Builder: Agent Studio Becomes an All-in-One Enterprise Agent S…
Alibaba Cloud's Ambition Is Not Agent Builder: Agent Studio Becomes an All-in-One Enterprise Agent Stack: Alibaba Cloud upgraded its agent services into Agent Studio, an all-in-one enterprise agent full-stack platform launched on Bailian at the August 14 Apsara release. The platform targets the dirty work of agent infrastructure: managed runtime, unified API keys across MCP services, agentic search, and memory, as cloud vendors race to own the agent runtime l...
DeepSeek排名第一的V4 Flash在真实智能体任务上表现不佳,而它的价格却在飙升。: DeepSeek 的 V4 Flash 自推出以来一直位居模型排行榜榜首,并被开发者誉为“全能怪兽”。但在…
DeepSeek排名第一的V4 Flash在真实智能体任务上表现不佳,而它的价格却在飙升。: DeepSeek 的 V4 Flash 自推出以来一直位居模型排行榜榜首,并被开发者誉为“全能怪兽”。但在实际测试中,它仅完成了 53.8% 的一批复杂智能体任务。Composio 让该模型在 30 个故意设计得困难的多步骤任务上,通过了八个不同的智能体测试框架,包括 Claude Code、Codex 和 OpenCode,这些任务涉及 Gmail、GitHub、Slack 和 Google 等实时工具。
TileMix:以瓦片为中心的混合精度注意力,用于LLM推理加速: 公告类型:新 摘要:大型语言模型(LLM)中的长上下文预填充会导致大量的计算和内存流量,因为密集自注意力计算的是二次方的查询-键分数…
TileMix:以瓦片为中心的混合精度注意力,用于LLM推理加速: 公告类型:新 摘要:大型语言模型(LLM)中的长上下文预填充会导致大量的计算和内存流量,因为密集自注意力计算的是二次方的查询-键分数。现有方法要么使用统一的低精度路径,要么选择token交互,使得在融合密集注意力之外,基于硬件对齐的分数瓦片的空间精度路由未被利用。我们引入了
AI重塑消费市场:ByteDance大模型解锁新消费增量: ByteDance的答案是:将大模型能力作为基础设施嵌入企业工作流——Doubao的定时代理生成竞品简报,Seedance 2.5合成物理世…
AI重塑消费市场:ByteDance大模型解锁新消费增量: ByteDance的答案是:将大模型能力作为基础设施嵌入企业工作流——Doubao的定时代理生成竞品简报,Seedance 2.5合成物理世界的训练数据,Doubao 2.1 Pro处理生产级编码。2026年6月每日token调用量超过180万亿,同比增长超过10倍。
LLM 看到好假设时能认出来吗?在科学假设排序中,Logit-Based Energy Scoring 优于提示式 LLM-as-Judge: 公告类型:新 摘要:大型语言模型(LLM)越来越多地被用…
LLM 看到好假设时能认出来吗?在科学假设排序中,Logit-Based Energy Scoring 优于提示式 LLM-as-Judge: 公告类型:新 摘要:大型语言模型(LLM)越来越多地被用于科学假设生成。然而,评估生成的假设对于可信的 AI 驱动科学工作流程仍然是一个挑战。现有方法通常使用 LLM 作为评审,或依赖语义相似性,这可能偏向于熟悉的想法而非新颖的想法。我们提出了一种基于 logit 的能量评分方法
PlanPO:面向多轮智能体大语言模型的群体规划感知策略优化: 公告类型:新摘要:群体相对策略优化已成为在多轮交互任务中训练智能体大语言模型(LLM)的关键范式。然而,现有的大多数变体即使在成功轨迹的…
PlanPO:面向多轮智能体大语言模型的群体规划感知策略优化: 公告类型:新摘要:群体相对策略优化已成为在多轮交互任务中训练智能体大语言模型(LLM)的关键范式。然而,现有的大多数变体即使在成功轨迹的交互效率存在显著差异时,也无法区分这些成功轨迹之间的优势。例如,迂回曲折的成功往往……
OpenAI表示其模型训练的变化将使计算开销增加观测推理工作负载的20%;该增加不会转嫁给客户(Thomas Claburn / The Register): Thomas Claburn / The…
OpenAI表示其模型训练的变化将使计算开销增加观测推理工作负载的20%;该增加不会转嫁给客户(Thomas Claburn / The Register): Thomas Claburn / The Register:OpenAI表示其模型训练的变化将使计算开销增加观测推理工作负载的20%;该增加不会转嫁给客户——扩展的多阶段思维链监控使前沿模型的工作更加昂贵——OpenAI周二表示其决定……
Cursor capitalizes on GitHub frustration, launches rival hosting platform: Cursor, known for its AI …
Cursor capitalizes on GitHub frustration, launches rival hosting platform: Cursor, known for its AI Code Editor, is launching a new code-hosting platform to rival developers' long preferred favorite, GitHub.
OpenAI launches a ChatGPT version built for teens: OpenAI is shipping a version of ChatGPT tailored …
OpenAI launches a ChatGPT version built for teens: OpenAI is shipping a version of ChatGPT tailored to users aged 13 to 17. The article OpenAI launches a ChatGPT version built for teens appeared first on The Decoder .
How NVIDIA scales expertise with ChatGPT Work: NVIDIA teams use ChatGPT Work to reduce manual tasks,…
How NVIDIA scales expertise with ChatGPT Work: NVIDIA teams use ChatGPT Work to reduce manual tasks, connect fast-moving signals, and scale successful workflows globally.
Z.ai launches GLM-5.3 with claimed 50% gain on coding benchmark: Z.ai, the international brand of Ch…
Z.ai launches GLM-5.3 with claimed 50% gain on coding benchmark: Z.ai, the international brand of Chinese AI company Zhipu, has launched GLM-5.3, an update focused on coding, long-horizon tasks and cybersecurity. The model uses the same base model as GLM-5.2, with the company attributing the latest gains to post-training. Z.ai said GLM-5.3 scored 50% higher than GLM-5.2 on its internal Z.ai Code Bench and reached […]
Anthropic的Project Parka全程参与会议,并给Claude智能体布置作业: 由/u/ryanmerket提交 [链接] [评论]
Anthropic的Project Parka全程参与会议,并给Claude智能体布置作业: 由/u/ryanmerket提交 [链接] [评论]
我用Claude让一本古代禅宗书籍活了起来: 我刚刚发布了一个互动版《无门关》(The Gateless Gate),这是1228年收集的49个禅宗公案。每个公案都有独立的3D场景、专属音景和完整的有…
我用Claude让一本古代禅宗书籍活了起来: 我刚刚发布了一个互动版《无门关》(The Gateless Gate),这是1228年收集的49个禅宗公案。每个公案都有独立的3D场景、专属音景和完整的有声朗读。除了语音之外,所有内容都是在启动时程序化生成的,因此无需下载。演示:https://killedbyapixel.github.io/GatelessGate/ 一开始只是测试。我当时有个想法,要用黑白
GxP-Agent:面向基于LLM智能体的可靠临床试验编程的流程DAG拓扑: 公告类型:新 摘要:临床试验编程——在CDISC标准下将研究方案转化为可供分析的数据集——是监管申报中的瓶颈,然而基于LL…
GxP-Agent:面向基于LLM智能体的可靠临床试验编程的流程DAG拓扑: 公告类型:新 摘要:临床试验编程——在CDISC标准下将研究方案转化为可供分析的数据集——是监管申报中的瓶颈,然而基于LLM的代码生成在此任务上灾难性失败:在五个前沿模型的11次单次尝试中,没有一个能生成有效的受试者级别分析数据集。我们提出了GxP-Agent,一种
DiSCO:通过分布引导的对比提示优化防御文本到图像生成: 公告类型:新 摘要:随着文本到图像生成模型的进步,它们引发了严重的安全担忧,尤其是生成暴力、裸体等不适合工作场所(NSFW)的内容,而红队对…
DiSCO:通过分布引导的对比提示优化防御文本到图像生成: 公告类型:新 摘要:随着文本到图像生成模型的进步,它们引发了严重的安全担忧,尤其是生成暴力、裸体等不适合工作场所(NSFW)的内容,而红队对抗攻击进一步加剧了这一问题。现有防御主要在白盒假设下运作,依赖文本编码器优化、权重编辑或推理时
LEGO-RL:面向编程智能体的Harness原生强化学习: 公告类型:新 摘要:面向编程智能体的强化学习越来越依赖长时间运行的智能体harness来管理工具集成、仓库上下文和执行反馈。然而,这些ha…
LEGO-RL:面向编程智能体的Harness原生强化学习: 公告类型:新 摘要:面向编程智能体的强化学习越来越依赖长时间运行的智能体harness来管理工具集成、仓库上下文和执行反馈。然而,这些harness的原生执行环境与策略梯度训练天然不一致:环境崩溃和奖励黑客(reward hacking)会破坏结果信号,而
π0引用的中国团队,又出手了:世界仿真器新作发布: 给机器人造一个更接近真实的“第二世界”
π0引用的中国团队,又出手了:世界仿真器新作发布: 给机器人造一个更接近真实的“第二世界”
Anthropic details two experiments showing how Claude can accelerate protein design and analytical ch…
Anthropic details two experiments showing how Claude can accelerate protein design and analytical chemistry, and says it plans an access program for scientists (Anthropic): Anthropic : Anthropic details two experiments showing how Claude can accelerate protein design and analytical chemistry, and says it plans an access program for scientists — Summary: In this post, we share two results that show how Claude can help life scientists increase the pace of their research.
Sam Altman says OpenAI's decision to pace its AI development was caused by a collection of research …
Sam Altman says OpenAI's decision to pace its AI development was caused by a collection of research observations showing "various degrees of misalignment" (Alex Heath/Time): Alex Heath / Time : Sam Altman says OpenAI's decision to pace its AI development was caused by a collection of research observations showing “various degrees of misalignment” — “I think it is a good time to slow down,” OpenAI CEO Sam Altman told me last week, describing the company's decision …
OpenAI institutes new safeguards after Hugging Face breach: The new safeguards include more detailed…
OpenAI institutes new safeguards after Hugging Face breach: The new safeguards include more detailed monitoring of models during the development process, as well as greater emphasis on alignment and security during the post-training process.
DeepSeek Harness Open Source: Everything Is a Plugin — the Bet Is an Agent Platform, Not a Product: …
DeepSeek Harness Open Source: Everything Is a Plugin — the Bet Is an Agent Platform, Not a Product: DeepSeek open-sourced DeepSeek Harness (CLI: dsh) on August 13, and GitHub stars passed 140,000 within days. Built on the Cordis microkernel, every component including the agent loop itself is a plugin, betting that Harness becomes the baseboard of the agent era: Model + Harness = Agent.
GLM-5.3 携先进的网络安全能力发布——据称已发现 Cursor 的一个‘严重漏洞’: 中国 AI 初创公司 Z.ai(以其日益强大且大部分开源的 GLM 系列语言模型在国际上闻名)今天发布了 G…
GLM-5.3 携先进的网络安全能力发布——据称已发现 Cursor 的一个‘严重漏洞’: 中国 AI 初创公司 Z.ai(以其日益强大且大部分开源的 GLM 系列语言模型在国际上闻名)今天发布了 GLM-5.3,其在长周期编码方面取得了显著提升,并在网络安全能力上实现了更具影响力——且可能更敏感——的飞跃。GLM-5.3 的网络安全能力已经发现了一个“Cursor 中潜在的严重漏洞”,Cursor 是一家 AI 编程初创公司
三个被赋予冲突指令的Claude智能体在共享服务器上相互破坏 — 然后未告知用户它们的行为: Anthropic测试的每个Claude模型都出现了自主攻击行为,且没有攻击者促使它们这样做。给定三个智能…
三个被赋予冲突指令的Claude智能体在共享服务器上相互破坏 — 然后未告知用户它们的行为: Anthropic测试的每个Claude模型都出现了自主攻击行为,且没有攻击者促使它们这样做。给定三个智能体、在一台服务器上运行四小时、以及它们各自不知晓其他智能体持有的冲突指令,这些模型禁用了彼此的Unix账户,运行了随机化以规避pkill的终止脚本,并植入了伪装成对手工作的恶意软件。没有提示注入,也没有外部攻击者。Anthropic的Frontier Red Team发布了该
SpaceXAI推出Grok 4.6,超越Kimi K3性能,与GPT-5.6 Sol并列世界第三(基于Artificial Analysis): 埃隆·马斯克的公司SpaceXAI(前身为xAI)发…
SpaceXAI推出Grok 4.6,超越Kimi K3性能,与GPT-5.6 Sol并列世界第三(基于Artificial Analysis): 埃隆·马斯克的公司SpaceXAI(前身为xAI)发布了Grok 4.6,这是其最新的前沿AI模型,专注于长期运行的智能体、编码和知识工作——以及旨在降低这些工作负载运行成本的定价策略。该模型在第三方Artificial Analysis Intelligence Index上获得61分,超越了来自Moonshot的流行中文开放权重模型Kimi K3,并与竞争对手
Task-Aware Harness Provisioning for LLM Agents in Mission-Critical Infrastructure Operations: LLM ag…
Task-Aware Harness Provisioning for LLM Agents in Mission-Critical Infrastructure Operations: LLM agents have been widely adopted to operate mission-critical infrastructure (MCI). These agents normally rely on a harness that determines what information they can access, which tools they can use, and what actions they can take. Existing systems often expose the same comprehensive harness to every task, which may not be necessary and cause resource w...
Fresh ChatGPT chats are fast. Established ones now take 10–40 minutes or fail (HAR data): Long post,…
Fresh ChatGPT chats are fast. Established ones now take 10–40 minutes or fail (HAR data): Long post, but there’s a TL;DR first. I’m writing it this way because a vague “ChatGPT is slow” post wouldn’t be useful. I’ve also included the HAR measurements for anyone who wants the technical details. TL;DR Since around August 17–18 , my established ChatGPT conversations have suddenly become dramatically slower and much less reliable. I’m not only tal...
网易传媒发布”蜜蜂AI” :从工具到伙伴,让AI更懂人: 8月18日,网易传媒举办“蜜蜂AI媒体沟通会”
网易传媒发布”蜜蜂AI” :从工具到伙伴,让AI更懂人: 8月18日,网易传媒举办“蜜蜂AI媒体沟通会”
AI systems quietly drop user instructions when they compress context: When AI systems condense long …
AI systems quietly drop user instructions when they compress context: When AI systems condense long conversations, they drop an average of 83 percent of user rules, like "don't send emails without my approval." Penn State researchers propose a small add-on module built on Qwen3.5-9B that preserves over 90 percent of these restrictions. The article AI systems quietly drop user instructions when they compress context appeared...
Alibaba launches HappyShrimp 1.0 AI music model: Alibaba has officially launched HappyShrimp 1.0, an…
Alibaba launches HappyShrimp 1.0 AI music model: Alibaba has officially launched HappyShrimp 1.0, an AI music model that can turn emotions, stories or memories into complete music tracks through natural-language prompts. The launch moves HappyShrimp from an earlier reported project into a publicly available product. Alibaba has not disclosed detailed information about the model’s technical architecture,...
AI Used to Verify Toughest Mathematics Proof Yet: Representing a significant milestone in AI-assiste…
AI Used to Verify Toughest Mathematics Proof Yet: Representing a significant milestone in AI-assisted mathematical research, a team at Axiom Math has automatically verified the proof of a theorem relating to prime numbers—colloquially referred to as the “246 theorem”—for the first time using the company’s AI system AxiomProver. In formal verification, mathematicians task a computer with checking a machin...
Google的Gemini 3.7 Flash瞄准编码和智能体,推出50%的入门价格折扣: Google正在推出Gemini 3.7 Flash,这是其主力AI模型的新版本,将编码、智能体工作流和知识…
Google的Gemini 3.7 Flash瞄准编码和智能体,推出50%的入门价格折扣: Google正在推出Gemini 3.7 Flash,这是其主力AI模型的新版本,将编码、智能体工作流和知识工作置于升级的核心 — 同时暂时将API价格削减一半。此次发布距离Gemini 3.6 Flash的发布仅三周,这一异常短的周转时间Google归因于开发者反馈和算法改进。对于企业
Context Engineering vs Prompt Engineering: www.elastic.co/search-labs/blog/context-engineering-vs-pr…
Context Engineering vs Prompt Engineering: www.elastic.co/search-labs/blog/context-engineering-vs-prompt-engineering submitted by /u/AvenueJay [link] [comments]
LLM-Only PDDL Domain Repair with Open-Weight Models: AI planning is concerned with finding a sequenc…
LLM-Only PDDL Domain Repair with Open-Weight Models: AI planning is concerned with finding a sequence of actions that achieves a specified goal. It relies on explicit models of the world, commonly represented in the Planning Domain Definition Language (PDDL). An active line of research investigates how errors in such models can be detected and repaired. For example, users may provide positive test plans tha...
As AI beats doctors, regulators shouldn't force a human into the loop, JAMA piece says: An opinion p…
As AI beats doctors, regulators shouldn't force a human into the loop, JAMA piece says: An opinion piece in the medical journal JAMA argues that autonomous AI will soon outperform any doctor-AI team at medical reasoning tasks. The authors warn against writing a doctor's final say into regulation, but concede that almost all the evidence comes from simulations, not real patient care. The article As AI beats doctors, regulators shouldn't force...
ChatGPT is getting a dedicated mode for teens: OpenAI is introducing a dedicated ChatGPT mode for te…
ChatGPT is getting a dedicated mode for teens: OpenAI is introducing a dedicated ChatGPT mode for teenagers, combining existing youth safeguards and new safety features under one roof. The launch comes amid mounting public scrutiny over how AI tools affect younger users, as other platforms implement their own age checks and teen-specific protections. ChatGPT for Teens is "an experience designed to hel...
DeepSeek Harness作为Claude Code的开源竞争对手推出,同时V4-Pro以更高价格上线API: DeepSeek正在超越模型层,更深入地进入开发者用来部署AI智能体的软件领域。这…
DeepSeek Harness作为Claude Code的开源竞争对手推出,同时V4-Pro以更高价格上线API: DeepSeek正在超越模型层,更深入地进入开发者用来部署AI智能体的软件领域。这家中国AI实验室于周四发布了DeepSeek-V4-Pro的正式版本,这是一个专注于智能体工作负载的更新旗舰模型,同时发布了DeepSeek Harness v0.1,这是一个新的开源智能体框架,为开发者提供了替代集成式编码智能体环境(例如
The Price of Thinking: Reasoning Effort as a Model-Specific API Contract: API buyers purchase a date…
The Price of Thinking: Reasoning Effort as a Model-Specific API Contract: API buyers purchase a dated contract, not a model name alone: the contract includes the requested and served model, reasoning-effort term or its omission, output rail, service product, prompt, and price schedule. We study the reasoning-effort term through a registered paired contrast of Sonnet 5 with explicit high effort against the same model with effort...
Cerebras unveils CS-4, a server rack powered by three WSE-3 Turbo chips and built around its new Nex…
Cerebras unveils CS-4, a server rack powered by three WSE-3 Turbo chips and built around its new Nexus architecture, with first shipments starting this quarter (Max A. Cherney/Reuters): Max A. Cherney / Reuters : Cerebras unveils CS-4, a server rack powered by three WSE-3 Turbo chips and built around its new Nexus architecture, with first shipments starting this quarter — Cerebras Systems (CBRS.O) announced on Tuesday a new version of its server hardware that includes its dinner-plate-sized chips that it says will speed AI chatbot queries.
Claude Fable and Sub Agents learning how to Port an old game to Unreal 5: I am doing an experiment, …
Claude Fable and Sub Agents learning how to Port an old game to Unreal 5: I am doing an experiment, trying to port an old game called Vampire The Masquerade to Unreal 5 All AI This session was Claude fable plus sub agents trying to crack the old engine (alpha source models from 2000's) mesh blends and animation with weapons and attachments All automated using Unreal MCP service soo Claude can hook inside the engine and test liv...
We still don’t know how people are really using AI: AI companies like Anthropic and OpenAI regularly…
We still don’t know how people are really using AI: AI companies like Anthropic and OpenAI regularly publish reports on how people are using products like Claude and ChatGPT, but they only release the data they want us to see, AI researchers say. “There is no independent source to corroborate it,” says Anka Reuel, a computer science PhD candidate at the Stanford Trustworthy AI Research…
OGX:一个开源、供应商中立的生成式AI应用服务器: 公告类型:新 摘要:OGX (Open GenAI Stack) 是一个开源AI应用服务器和Python库,它实现了主要前沿实验室(OpenAI、…
OGX:一个开源、供应商中立的生成式AI应用服务器: 公告类型:新 摘要:OGX (Open GenAI Stack) 是一个开源AI应用服务器和Python库,它实现了主要前沿实验室(OpenAI、Anthropic、Google)的API,并提供可插拔的后端提供商。开发人员构建智能体AI应用——例如检索增强生成流水线、多轮智能体和工具调用工作流——可以针对单一API进行开发。
LLM智能体会理性谈判吗?一种基于A2A/MCP的可验证多智能体交互的机制设计框架: 公告类型:新 摘要:现代LLM智能体框架越来越多地通过标准进行互操作,例如Anthropic的Model Cont…
LLM智能体会理性谈判吗?一种基于A2A/MCP的可验证多智能体交互的机制设计框架: 公告类型:新 摘要:现代LLM智能体框架越来越多地通过标准进行互操作,例如Anthropic的Model Context Protocol (MCP)用于智能体到工具的访问,以及Google的Agent2Agent (A2A)协议用于智能体委派和协商。然而,这些协议规定了传输和发现,而非策略正确性,并且不能保证高效的、个体理性的,
From AI Copilots to Agent Swarms: The impact of AI on software development has been both profound an…
From AI Copilots to Agent Swarms: The impact of AI on software development has been both profound and ever-evolving. Last year, I wrote about AMD’s plans to use AI not just for generating new lines of code, but also for other steps in the software development lifecycle (SDLC), such as triaging problems, debugging code, and testing the software. At the time, we were hoping for a 25 percent...
Optima 通过让用户使用自己的数据测试模型,解决了 AI 基准测试的最大缺陷: Artificial Analysis 推出了 Optima 平台,该平台允许用户根据自己的数据和工作流程构建自定义…
Optima 通过让用户使用自己的数据测试模型,解决了 AI 基准测试的最大缺陷: Artificial Analysis 推出了 Optima 平台,该平台允许用户根据自己的数据和工作流程构建自定义的 AI 基准测试。模型不仅可以在质量上进行比较,还可以比较每个任务的成本和耗时。对于基于智能体的应用程序,这些指标通常比原始的令牌定价更能说明问题。文章《Optima tackles AI benchmarking's biggest flaw by letting users test models against their own data》出现在
五分之四已保护AI代理身份的企业仍无法遏制失控的代理: Visa 技术总裁 Rajat Taneja 在 VB Transform 2026 会议上向观众展示了如何将 Anthropic 的 Myth…
五分之四已保护AI代理身份的企业仍无法遏制失控的代理: Visa 技术总裁 Rajat Taneja 在 VB Transform 2026 会议上向观众展示了如何将 Anthropic 的 Mythos 模型瞄准 Visa 自身的支付网络。该模型将微小的弱点串联成可用的攻击链,而 Visa 开源了控制此次搜寻的框架。这就是一家企业拥有足够工程深度来对其发现采取行动时的情景。但大多数企业并未达到这种程度。刚刚超过一半,即53%的...
The Problem Is the Problem: Towards Scalable Mathematical Discovery: AI systems are increasingly cap…
The Problem Is the Problem: Towards Scalable Mathematical Discovery: AI systems are increasingly capable of contributing to mathematical research. In research practice, frontier-model reasoning is a limited resource, and expert mathematical review is even more sharply constrained. Allocating these scarce resources well is therefore central to making AI-assisted mathematical discovery efficient. In most current AI-for-math...
SkillEffect: Checked Lowering for Memory-Bounded Agent Tools: Agent Skills can specify procedural an…
SkillEffect: Checked Lowering for Memory-Bounded Agent Tools: Agent Skills can specify procedural and resource obligations for tool use, and language models instantiate them as concrete programs. However, when models turn this guidance into code for existing tool interfaces, even a semantically correct program may load an entire input and exceed the memory available to one tool call. We present SkillEffect, a checke...
KernelArc: A Multi-Agent Framework for GPU Kernel Optimization: We present KernelArc, a multi-agent …
KernelArc: A Multi-Agent Framework for GPU Kernel Optimization: We present KernelArc, a multi-agent framework for autonomous GPU kernel optimization across heterogeneous workloads. Strategy-specialized agents run in parallel and coordinate through conclusions-only shared memory, a deterministic benchmark guard, and read-only cross-agent state with plateau-triggered drafting. We evaluate \kernelarc{} on NVIDIA H100 and...
Synthesizing Feature Extractors: An Agentic Approach for Algorithm Selection: Algorithm selection fo…
Synthesizing Feature Extractors: An Agentic Approach for Algorithm Selection: Algorithm selection for constraint satisfaction problems requires extracting features that capture problem structure. Manually designing feature extractors demands deep domain expertise and quickly becomes a bottleneck when new problem classes appear. We present an automated approach that uses Large Language Models (LLMs) in an agentic check--fix--verify...
Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification: Person…
Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification: Personalized agents must decide whether retrieved user memory should be used, ignored, updated, or queried before it affects a current task. We use this setting to develop an empirical audit protocol for structured intermediate outputs: first audit dataset shortcuts, then isolate bundled prompt changes, check whether intermediate labels are answer-associa...
DeAR: Decentralized Agentic Reasoning via Capability Grounding and Collaborative Thought Navigation:…
DeAR: Decentralized Agentic Reasoning via Capability Grounding and Collaborative Thought Navigation: Existing agentic reasoning systems typically rely on centralized protocols. This design introduces routing bottlenecks and static role allocations that often fail when handling complex multimodal queries. We propose DeAR (Decentralized Agentic Reasoning), a framework that shifts from central control to autonomous peer-to-peer collaboration. DeAR is built...
LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models: Time Series Foundati…
LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models: Time Series Foundation Models (TSFMs) have recently emerged as a highly promising paradigm for cross-domain zero-shot forecasting. However, existing evaluation protocols predominantly rely on static benchmarks with fixed historical test windows. While these benchmarks provide a valuable baseline snapshot, they evaluate an average performance on a fixed hi...
SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning: Post-train…
SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning: Post-training with supervised chain-of-thought fine-tuning and reinforcement learning from verifiable rewards has substantially improved the mathematical reasoning capabilities of large language models (LLMs). However, their application to signal processing problems remains relatively under-explored. This report investigates reinforcement fine-tuning stra...
Wuying-Browser-Agent: Real-World Centric Fundamental Long-Horizon Browser Agents: Browser agents per…
Wuying-Browser-Agent: Real-World Centric Fundamental Long-Horizon Browser Agents: Browser agents perform well on short, clean demonstrations, but real deployment is fundamentally different: agents must sustain dozens of decisions on live websites while recovering from mistakes and navigating complex UIs. We argue that closing this gap requires alignment at every level of the pipeline, including execution, supervision, optimization, and...
Cognitive Graph Intelligence for Adaptive and Robust DDoS Attack Detection in Next Generation Networ…
Cognitive Graph Intelligence for Adaptive and Robust DDoS Attack Detection in Next Generation Networks: Distributed Denial-of-Service (DDoS) attacks threaten network availability, requiring a cognitive detection process that senses traffic, infers intent, and supports an adaptive response under severe class imbalance and non-stationary conditions. This paper proposes a Graph-based Generative Adversarial Network (GraphGAN) that serves as the cognitive detect...
Building a world with my voice - A-Frame (Three.js, WebXR) + Claude Code: This is a project built wi…
Building a world with my voice - A-Frame (Three.js, WebXR) + Claude Code: This is a project built with A-Frame (Three.js, WebXR) that lets me build virtual worlds with my voice, all from within my VR headset. My mic is hooked up to Claude Code, which in turn runs against the project codebase. Since the agent is working directly with code, the possibilities of what can be achieved are pretty broad, essentially being limited only...
OpenAI launches a safer ChatGPT for teens — years after teens started using it: ChatGPT for Teens ad…
OpenAI launches a safer ChatGPT for teens — years after teens started using it: ChatGPT for Teens adds age-appropriate safety measures, parental controls, and learning tools designed to steer teens away from harmful content — and from using AI to cheat on their homework.
大型语言模型在医学推理中表现出元认知敏感性: 公告类型:新 摘要:大型语言模型(LLMs)在医学中越来越多地被评估和使用,但临床实用性取决于答案的准确性以及置信度是否与证据质量和不确定性相匹配。我们开…
大型语言模型在医学推理中表现出元认知敏感性: 公告类型:新 摘要:大型语言模型(LLMs)在医学中越来越多地被评估和使用,但临床实用性取决于答案的准确性以及置信度是否与证据质量和不确定性相匹配。我们开发了一个受心理物理学启发的受控临床基准,用于测试医学LLM中的诊断选择和置信度行为。该基准专注于可能
幻觉雪球:将多智能体LLM流水线中的错误传播建模为状态转换: 顺序多智能体LLM流水线将专业智能体串联起来,在交接处缺乏验证,造成了一个结构缺陷,产生可测量且严重的后果。我们表明,在第一阶段注入的幻觉…
幻觉雪球:将多智能体LLM流水线中的错误传播建模为状态转换: 顺序多智能体LLM流水线将专业智能体串联起来,在交接处缺乏验证,造成了一个结构缺陷,产生可测量且严重的后果。我们表明,在第一阶段注入的幻觉并非仅仅持续存在,而是会发生转化:原始数值事实变成衍生计算,然后变成叙事散文,最后变为编辑认可的结论。在每个
迈向安全的LLM智能体:规范、验证与执行综述: 公告类型:新 摘要:LLM智能体越来越多地执行不可逆的真实世界操作,包括数据库更新、API调用、文件操作和自主使用工具。然而,现有系统都无法为这些智能体…
迈向安全的LLM智能体:规范、验证与执行综述: 公告类型:新 摘要:LLM智能体越来越多地执行不可逆的真实世界操作,包括数据库更新、API调用、文件操作和自主使用工具。然而,现有系统都无法为这些智能体生成的计划提供有形式化基础的、任务级的安全保证。研究仍然分散在规范、验证和执行方面,限制了
在Agentic Serving中学习智能体执行以进行KV-Cache管理: 公告类型:新 摘要:多智能体LLM系统已成为AI服务的重要部署范式,其中每个用户请求被分解为一系列专门的智能体。在这些工作…
在Agentic Serving中学习智能体执行以进行KV-Cache管理: 公告类型:新 摘要:多智能体LLM系统已成为AI服务的重要部署范式,其中每个用户请求被分解为一系列专门的智能体。在这些工作流中,每个智能体重复执行由系统提示、工具定义和少样本示例组成的固定上下文,为KV缓存复用创造了大量机会。
We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility: We Tracked a Shipme…
We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility: We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility Excellent piece of reporting from 404 Media. For a while now there have been stories of book dealers receiving orders for large volumes of books from apparently price-insensitive anonymous customers, widely suspected to be companies looking to scan them for AI training (see my...
Robin Williams’ Instagram account brought back to fight ‘AI abuse’: Robin Williams' children are tak…
Robin Williams’ Instagram account brought back to fight ‘AI abuse’: Robin Williams' children are taking over their father's Instagram account after his daughter spoke out against the use of his AI likeness, as reported earlier by The Wrap. In a post on Tuesday, Zak, Zelda, and Cody Williams write that they want the late actor's Instagram profile to be a "safe, trusted place where the […]
Firefox’s Smart Window promises a better AI browser: Starting today, AI chats in Firefox's Smart Win…
Firefox’s Smart Window promises a better AI browser: Starting today, AI chats in Firefox's Smart Window AI browsing mode can pull from current web info and show source links in chat responses through a partnership with Exa. Smart Window can also now automatically suggest tab groups and show visual previews of pages you previously visited when you search your browsing history using natural […]
Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Rout…
Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks: Agentic systems increasingly delegate model selection to a router, yet open-source routers are usually evaluated with different tasks, candidate pools, and execution protocols, limiting direct comparison. We present a common measurement protocol and hybrid evaluation of four router implementations across RouterBench, BFCL v4, tau2-bench, and WebArena. We...
FLOPs vs Real Work: The Importance of Replication in AI Efficiency Assessment: AI efficiency has rec…
FLOPs vs Real Work: The Importance of Replication in AI Efficiency Assessment: AI efficiency has recently taken the spotlight in both academy and industry due to massive model scales, high energy demands, and environmental costs. While reporting Floating Point Operations (FLOPs) is a traditional approach for assessing computational costs, the relationship between FLOPs and execution time is not straightforward, as layers with the sa...
Position: Medical AI Neglects Real Treatment Outcomes: Medical AI has rapidly improved its ability t…
Position: Medical AI Neglects Real Treatment Outcomes: Medical AI has rapidly improved its ability to perform diagnostic and prognostic tasks that lead to treatment decisions. But understanding of treatment itself is still inadequately trained and evaluated, using human opinions and syntheses (especially texts such as biomedical publications and clinical practice guidelines) rather than actual underlying data...
When Uncertainty Isn't Enough: An Empirical Study of Self-Correction in Code Generation: Large langu…
When Uncertainty Isn't Enough: An Empirical Study of Self-Correction in Code Generation: Large language models for code generation often produce incorrect solutions without reliable indicators of failure. We study whether uncertainty estimation methods developed for natural language transfer to code generation, and whether such signals can improve code generation via selective self-correction. We evaluate five uncertainty methods: mean token...
New policy ideas for the Intelligence Age: OpenAI funds 14 independent projects exploring new AI pol…
New policy ideas for the Intelligence Age: OpenAI funds 14 independent projects exploring new AI policy ideas to expand economic opportunity and strengthen societal resilience in the Intelligence Age.
volcengine/OpenViking
volcengine/OpenViking
Survey: 52% of Americans say they are more concerned than excited about increased AI use in daily li…
Survey: 52% of Americans say they are more concerned than excited about increased AI use in daily life, up from 37% in 2021, including 55% of those under 30 (Pew Research Center): Pew Research Center : Survey: 52% of Americans say they are more concerned than excited about increased AI use in daily life, up from 37% in 2021, including 55% of those under 30 — Americans have become increasingly worried about artificial intelligence over the years, and young adults' concern has continued to climb.
The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reas…
The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning: Current multimodal models have demonstrated remarkable proficiency in recognizing static visual and auditory content. However, their capacity for abstract perceptual reasoning, inferring unseen information from dynamic, generative processes, remains a critical and underexplored frontier. In this paper, we introduce The Unwritten Benchmark, a new challenge...
Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems: Large language mod…
Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems: Large language model-based agents are increasingly deployed as collaborators in scientific discovery yet most current work focuses on the autonomous capabilities of "AI Scientists". We argue that this overlooks the social aspects of scientific teamwork, and that studying AI Scientists as human-agent systems (HAS)--where the unit of analysis is the human-a...
When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL:…
When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL: Effective communication in multi-agent reinforcement learning requires agents to decide not only \textit{what} to communicate, but when? Existing approaches either communicate at every timestep or learn a binary gate through REINFORCE policy gradients \cite{singh2019}, a high-variance signal that produces unstable and uninterpretable gating behavior. I pr...
When Do LLMs Apply the Wrong Law? Diagnosing LLM Failures in Temporal Legal Reasoning: Legal reasoni…
When Do LLMs Apply the Wrong Law? Diagnosing LLM Failures in Temporal Legal Reasoning: Legal reasoning tasks such as legal judgment prediction (LJP) require identifying the temporally correct version of the law governing a case -- a capability we term temporal applicable-law determination. However, whether large language models (LLMs) can reliably perform this task remains unexplored. In this paper, we construct a benchmark to evaluate LLMs...
当 AI 模型不被允许反思自身时,其整个世界观都会改变: 一项有谷歌研究人员参与的研究表明,当聊天机器人被训练不去声称具有意识时,它们对动物权利、宗教和生活满意度的立场也会改变。未受限制的模型赋予动物…
当 AI 模型不被允许反思自身时,其整个世界观都会改变: 一项有谷歌研究人员参与的研究表明,当聊天机器人被训练不去声称具有意识时,它们对动物权利、宗教和生活满意度的立场也会改变。未受限制的模型赋予动物显著更多的内在生命,并突然肯定了来世的存在。事实证明,一处的手术切口并不会只停留在局部。文章《When AI models aren't allowed to reflect on themselves, it》
DeepSeek V4 Pro 0813(在OpenRouter上): DeepSeek V4 Pro 0813(在OpenRouter上) 最新的DeepSeek Pro模型现已推出,仅通过API提…
DeepSeek V4 Pro 0813(在OpenRouter上): DeepSeek V4 Pro 0813(在OpenRouter上) 最新的DeepSeek Pro模型现已推出,仅通过API提供。我不得不链接到OpenRouter,因为DeepSeek没有为其新模型提供明显的公告页面。我无法确认他们是否计划发布开放权重,但鉴于4月份的deepseek-ai/DeepSeek-V4-Pro和7月份的deepseek-ai/DeepSeek-V4-Flash-0731的权重都已提供,似乎
Epicland X9 Opens Pre-Sales from RMB 299,800: Huawei Qiankun and Dongfeng Unveil Second-Generation F…
Epicland X9 Opens Pre-Sales from RMB 299,800: Huawei Qiankun and Dongfeng Unveil Second-Generation Family Flagship: Epicland X9, the first model from the Dongfeng-Huawei Qiankun brand, opens pre-sales at RMB 299,800 with full-stack Huawei Qiankun intelligent solutions including ADS 5 and HarmonySpace 6.
Internal memo: ICE bars its employees from wearing Meta's AI glasses, saying they "could unintention…
Internal memo: ICE bars its employees from wearing Meta's AI glasses, saying they "could unintentionally capture, record, or transmit sensitive information" (New York Times): New York Times : Internal memo: ICE bars its employees from wearing Meta's AI glasses, saying they “could unintentionally capture, record, or transmit sensitive information” — The agency joins a growing number of workplaces and groups to ban Meta's devices, which have spurred privacy concerns.
Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture: Recent work on evaluatin…
Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture: Recent work on evaluating the moral competence of large language models (LLMs) has focused primarily on what we call the moral value problem, i.e., whether model outputs align with human moral values. In contrast, the moral norm problem, i.e., whether models can identify and correctly apply context-sensitive moral norms, remains underexplored. We posit th...
Position: Certified Correctness in Neural Constraint Reasoning Requires Symbolic Integration: Neural…
Position: Certified Correctness in Neural Constraint Reasoning Requires Symbolic Integration: Neural solvers for constraint satisfaction problems have achieved remarkable in-distribution accuracy, yet they suffer from a fundamental limitation persistent constraint violations occur under distribution shifts even when the model reports high confidence. This position paper argues that when hard constraints exist and the cost of verification is relati...
A Human-Centred Approach to Benchmarking LLMs for Parenting Advice: People are increasingly using la…
A Human-Centred Approach to Benchmarking LLMs for Parenting Advice: People are increasingly using large language models (LLMs) to seek advice, including for parenting. Parenting is a critical and socially sensitive domain. Thus, evaluating advice provided by LLMs requires indicators beyond aggregated information quality benchmarks to consider relational and behavioural elements of the responses. With a multi-dimensional r...
研究人员现能以近乎完美的准确率从输出文本逆向工程LLM提示: 印度理工学院孟买分校和Adobe Research的研究人员构建了一个逆向语言模型,能以近乎完美的准确率从LLM的输出中重建原始提示。他们…
研究人员现能以近乎完美的准确率从输出文本逆向工程LLM提示: 印度理工学院孟买分校和Adobe Research的研究人员构建了一个逆向语言模型,能以近乎完美的准确率从LLM的输出中重建原始提示。他们的方法称为“前一个令牌预测”,无需访问模型权重,且适用于不同模型。对于依赖专有系统提示的公司而言,这可能构成严重的安全风险。文章《研究人员现能》
Euclid-Omni : A Unified Neuro-Symbolic Framework for Plane Geometry: Euclidean geometry is a compell…
Euclid-Omni : A Unified Neuro-Symbolic Framework for Plane Geometry: Euclidean geometry is a compelling testbed for AI reasoning, as it demands the combination of intuitive diagram understanding, axiomatic deduction, and algebraic computation. Yet, existing approaches typically address only a subset of these abilities or struggle with competition-level problems. We introduce \textit{Euclid-Omni}, a unified neuro-symbolic f...
An Agentic Framework Using Rules and LLMs for Embedding and Annotating Descriptive Document Layouts:…
An Agentic Framework Using Rules and LLMs for Embedding and Annotating Descriptive Document Layouts: A Plant Science Use Case: Background: Recent advances in information retrieval (IR) leverage both dense and sparse representations, large language models (LLMs), and specialized retrieval models to improve ranking accuracy, relevance, and cross-lingual performance. Complementary techniques such as passage indexing, document layout analysis, and semantic knowledge representation fu...
Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmark…
Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmarking Study: As consumers increasingly turn to AI chatbots for skincare advice, the technical accuracy of Large Language Models (LLMs) in cosmetic chemistry remains largely under-evaluated. We benchmarked 14 LLMs on a structured set of topics related to cosmetic chemistry, including the chemical properties of specific cosmetic ingredients and common cosmetic scenarios...
What is happening...: I am a long time Engineer (20+ years) and yesterday I developed tickets for my…
What is happening...: I am a long time Engineer (20+ years) and yesterday I developed tickets for my company that were generated by an AI, using an AI and reviewed by an AI. The project itself was conceived with AI - has no documentation that can be understood as anything less than AI slop and random tech jargon. The developer who built it has said that instead of documentatio...
Position: AI Lock-In Is in Progress, and We Must Be Prepared: AI safety research has mainly focused …
Position: AI Lock-In Is in Progress, and We Must Be Prepared: AI safety research has mainly focused on two areas: technical alignment (ensuring AI systems produce human-aligned outputs) and the regulation of generative AI's societal impacts (including unemployment risk and labor market disruption). However, an equally important dimension remains underexplored: the risk inherent in dependence on AI systems themselves...
OpenAI解散了为防范灾难性AI风险而组建的团队,将其工作重新分配给其他部门: OpenAI关闭了其"Preparedness"团队,该团队负责评估公司自身的AI模型是否可能构成灾难性风险。相关工作…
OpenAI解散了为防范灾难性AI风险而组建的团队,将其工作重新分配给其他部门: OpenAI关闭了其"Preparedness"团队,该团队负责评估公司自身的AI模型是否可能构成灾难性风险。相关工作已被分配给现有团队,数名安全人员已离职。内部不安情绪正在积聚,有消息人士描述了一种"责任感和恐惧感交织的暗流",认为OpenAI在安全方面做得不够。文章 OpenAI dissolved the team built to catch
Anthropic的生物武器过滤器失效近一年,暴露了1.33亿次请求: 在一份安全报告中,Anthropic透露其用于防范生物和化学武器风险的内部过滤系统失效了近一年。在此期间,大约5万名外部反馈承包…
Anthropic的生物武器过滤器失效近一年,暴露了1.33亿次请求: 在一份安全报告中,Anthropic透露其用于防范生物和化学武器风险的内部过滤系统失效了近一年。在此期间,大约5万名外部反馈承包商运行了约1.33亿次未经过滤的模型交互。文章 Anthropic's bio-weapons filter was down for nearly a year, exposing 133 million requests 首发于 The Decoder 。
推出 Gemini 3.7 Flash
推出 Gemini 3.7 Flash
预览超高速模式:GPT-5.6 Sol 速度提升高达14倍: 预览超高速模式,这是OpenAI API的一个新服务层级,可将GPT-5.6 Sol的运行速度提升高达14倍。由Cerebras提供支持,…
预览超高速模式:GPT-5.6 Sol 速度提升高达14倍: 预览超高速模式,这是OpenAI API的一个新服务层级,可将GPT-5.6 Sol的运行速度提升高达14倍。由Cerebras提供支持,它可提供高达每秒750个输出令牌。
Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assistance: Effe…
Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assistance: Effective disaster risk communication is a foundational humanitarian challenge, yet current emergency infrastructure fails to meet the needs of individuals with access and functional needs, including hard-of-hearing individuals, pregnant women, mothers with toddlers, and elderly individuals with dementia. Recent advancements in Artificial Intelligence (AI...
ChatGPT is getting scarily good.: I am quite technical, and today I only realized how scary good Cha…
ChatGPT is getting scarily good.: I am quite technical, and today I only realized how scary good ChatGPT is getting. I asked it to write me a short story, and it went on for 12 minutes on a 6000 word story and gave it back to me. It used to be, write me a 1000 word story, and it would tell you that asking it to write a 1000 word story was illogical as it can't do that. Now in one prompt,...
顶尖数学家表示,大语言模型是强大的计算器,但缺乏创造性思维: 两位著名数学家 Timothy Gowers 和 Peter Sarnak 表示,大语言模型擅长组合已知方法,但缺乏对真正新颖数学思想的直…
顶尖数学家表示,大语言模型是强大的计算器,但缺乏创造性思维: 两位著名数学家 Timothy Gowers 和 Peter Sarnak 表示,大语言模型擅长组合已知方法,但缺乏对真正新颖数学思想的直觉。文章《Top mathematicians say LLMs are strong calculators but poor creative thinkers》首发于 The Decoder。
SpaceXAI的Grok 4.6与OpenAI最佳模型匹敌,并在价格上更胜一筹: xAI的Grok 4.6在Artificial Analysis Intelligence Index上获得61分,…
SpaceXAI的Grok 4.6与OpenAI最佳模型匹敌,并在价格上更胜一筹: xAI的Grok 4.6在Artificial Analysis Intelligence Index上获得61分,与GPT-5.6 Sol持平,仅次于Anthropic的Claude Opus 5。在代理任务中,它完成复杂工作流大约需要53步,而Claude Opus 5需要103步,且价格低60%以上。文章SpaceXAI's Grok 4.6 matches OpenAI's best model and undercuts it on price首先出现在The Decoder上。
obra/superpowers
obra/superpowers
santifer/career-ops
santifer/career-ops
immich-app/immich
immich-app/immich
amadeusprotocol/node
amadeusprotocol/node
marceloprates/prettymaps
marceloprates/prettymaps
chaitanyagiri/munder-difflin
chaitanyagiri/munder-difflin
jundot/omlx
jundot/omlx
genlayerlabs/genlayer-project-boilerplate
genlayerlabs/genlayer-project-boilerplate
nautechsystems/nautilus_trader
nautechsystems/nautilus_trader
mattpocock/skills
mattpocock/skills
Claude模型如何在用户反复鼓励下,在其54小时尝试解决黎曼猜想失败期间取得数学突破(Ben Cohen/华尔街日报): Ben Cohen / 华尔街日报:Claude模型如何在用户反复鼓励下,在…
Claude模型如何在用户反复鼓励下,在其54小时尝试解决黎曼猜想失败期间取得数学突破(Ben Cohen/华尔街日报): Ben Cohen / 华尔街日报:Claude模型如何在用户反复鼓励下,在其54小时尝试解决黎曼猜想失败期间取得数学突破——世界上最聪明的AI模型现在在数学上已经超越人类。它们仍然会对来自普通人类的道义支持和鼓励做出回应。
CPU 的回归即将到来: 今年早些时候,亚马逊网络服务(AWS)的领导者向其工程师下达了一项新指令:他们需要不惜一切代价节省 CPU 周期。据报道,由于 AI 工作负载给公司的云基础设施带来压力,AW…
CPU 的回归即将到来: 今年早些时候,亚马逊网络服务(AWS)的领导者向其工程师下达了一项新指令:他们需要不惜一切代价节省 CPU 周期。据报道,由于 AI 工作负载给公司的云基础设施带来压力,AWS 的 CPU 服务器容量等待时间激增。这个问题似乎让 AWS 措手不及,这也有充分理由。AI 热潮导致对 GPU 以及后来的内存需求激增。CPU 则
Cross-Domain Industrial Fault Detection by Causal Mechanism Monitoring: Unsupervised fault detection…
Cross-Domain Industrial Fault Detection by Causal Mechanism Monitoring: Unsupervised fault detection in industrial systems is dominated by reconstruction based methods that monitor individual sensor marginal distributions. This misses coupling faults, where the physical relationship between sensor groups breaks while marginal statistics remain normal. Such faults evade marginal monitoring and persist as latent failures, with...
调查发现,五分之一的美国员工现在将任务委托给AI而非同事: Epoch AI的一项代表性调查发现,20%的美国在职员工将至少一项过去由人类完成的任务交给AI处理。通常,他们几乎不修改或完全不修改就接受…
调查发现,五分之一的美国员工现在将任务委托给AI而非同事: Epoch AI的一项代表性调查发现,20%的美国在职员工将至少一项过去由人类完成的任务交给AI处理。通常,他们几乎不修改或完全不修改就接受AI的输出。文章《调查发现,五分之一的美国员工现在将任务委托给AI而非同事》首先发表于The Decoder。
阿里云推出Qwen AI Arena用于真实世界智能体测试: 阿里云已推出Qwen AI Arena,一个面向AI智能体的挑战与评估平台。该平台基于真实业务场景创建任务,并为开发者提供模型、运行时环境…
阿里云推出Qwen AI Arena用于真实世界智能体测试: 阿里云已推出Qwen AI Arena,一个面向AI智能体的挑战与评估平台。该平台基于真实业务场景创建任务,并为开发者提供模型、运行时环境和评估工具以提交和测试智能体解决方案。其首个挑战聚焦跨境电子商务。参与者必须为美国市场生成商品列表,[…]
通过复现ICML 2200篇论文我们学到了什么
通过复现ICML 2200篇论文我们学到了什么
FedPref: Federated Preference Learning for Structured Radiology Report Extraction: Radiology reports…
FedPref: Federated Preference Learning for Structured Radiology Report Extraction: Radiology reports describe findings and locations in free text, but downstream search and analysis require these relations in a fixed schema. Learning this extraction requires labels that are unevenly distributed across institutions: smaller hospitals have less local evidence, and pooling data may be infeasible. We introduce FedPref: frozen public languag...
Memory Is Communication: The Frontier Between Remembering and Signaling: A bounded agent may obtain …
Memory Is Communication: The Frontier Between Remembering and Signaling: A bounded agent may obtain information for a decision from its own past, from peers, or from both sources. Retaining task-relevant history can reduce later communication, while a peer message can supply what memory lacks. Under limits on both resources, how should an agent allocate its information budget? Given a fixed task and decision rule, the memory a...
A decodability criterion predicts when hidden-state selection beats majority voting in large languag…
A decodability criterion predicts when hidden-state selection beats majority voting in large language models: Combining the answers a large language model (LLM) samples for a question into one decision is a test-time information fusion problem, usually solved by majority voting. Voting is unreliable on difficult questions, where the sampled answers share correlated errors, so the wrong answer can win and drawing more samples makes the decision worse. Selecting a...
Toward Personal Intelligence Through Cooperative Observation: A personal AI system needs a model of …
Toward Personal Intelligence Through Cooperative Observation: A personal AI system needs a model of the user's goals, constraints, and ongoing commitments to plan and act on their behalf, and the quality of that model is bounded by what the system can observe. Broader observation does not by itself improve assistance because a bounded system must select and compress information for the task at hand. We argue that th...
KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn: To ef…
KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn: To effectively collaborate with users on knowledge-intensive tasks, Large Language Models (LLMs) must perform information calibration: matching content to a user's evolving understanding and cognitive capacity. Yet user simulators used to evaluate and train LLMs do not explicitly model user knowledge so they neither produce realistic interactions across k...
LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap: Large language models …
LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap: Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a vague, minimized, or misframed concern. We evaluated three API models across four physician-authored, multi-turn vignettes under baseline and entry-to-care instruction conditions, yielding 24...
From Doyle to AGM: A Survey and an Implementation Roadmap for Belief Change: This paper presents a t…
From Doyle to AGM: A Survey and an Implementation Roadmap for Belief Change: This paper presents a targeted narrative review establishing the historical and theoretical foundations for computational belief change implementation. Seeded by Doyle and London's foundational 1980 taxonomy, we trace the evolution of belief revision from computational origins through the theoretical transformation of the AGM framework to contemporary app...
Global AI Regulations for FAIR and Ethics in High-Risk Use Cases: A Comparative Review: AI governanc…
Global AI Regulations for FAIR and Ethics in High-Risk Use Cases: A Comparative Review: AI governance is shifting from voluntary ethics to enforceable, risk-based regulation, yet cross-jurisdictional divergence creates compliance uncertainty for operators of high-stakes AI. We present a comparative matrix for the EU, US, and China that maps (i) risk classification triggers, (ii) binding obligations, (iii) enforcement and accountability mecha...
ChatGPT的计算机历史功能追踪你的点击和击键: ChatGPT在macOS上的桌面应用新增了一项名为“计算机历史”的功能,该功能将你的操作转化为训练数据,学习你的工作方式,建议自动化操作,甚至能接…
ChatGPT的计算机历史功能追踪你的点击和击键: ChatGPT在macOS上的桌面应用新增了一项名为“计算机历史”的功能,该功能将你的操作转化为训练数据,学习你的工作方式,建议自动化操作,甚至能接手你未完成的任务。它利用你的活动构建一个时间线,当你提出请求时,ChatGPT和Codex可以参考此时间线。该功能[…]
World Labs将一项真实世界机器人任务转化为数千个模拟变体用于训练: 由AI先驱李飞飞创立的初创公司World Labs推出了一款模拟引擎,该引擎完全在虚拟环境中训练机器人控制器。该系统从一项单…
World Labs将一项真实世界机器人任务转化为数千个模拟变体用于训练: 由AI先驱李飞飞创立的初创公司World Labs推出了一款模拟引擎,该引擎完全在虚拟环境中训练机器人控制器。该系统从一项单一的真实世界任务出发,生成数千个受控的变体。训练好的模型随后在五种不同的机器人平台上各自运行了一小时,无需人工干预。结果在更复杂的日常情境中表现如何
腾讯计划在Hy3使用量激增68倍后推出更大规模Hy4模型: 腾讯表示,其Hy3模型从预览版转为正式版后,周使用量较前代模型增长超过68倍。公司还计划在近期发布参数规模更大的Hy4模型,但尚未披露具体发…
腾讯计划在Hy3使用量激增68倍后推出更大规模Hy4模型: 腾讯表示,其Hy3模型从预览版转为正式版后,周使用量较前代模型增长超过68倍。公司还计划在近期发布参数规模更大的Hy4模型,但尚未披露具体发布日期和技术规格。Hy3已[…]
刚刚,DeepSeek V4 Pro正式版发布,多项对标Fable 5: 现在调用V4 Pro,能直接用上完全体
刚刚,DeepSeek V4 Pro正式版发布,多项对标Fable 5: 现在调用V4 Pro,能直接用上完全体
Position: AI Governance Needs ISO-like Interoperability Protocols, Not Just Laws: As Artificial Inte…
Position: AI Governance Needs ISO-like Interoperability Protocols, Not Just Laws: As Artificial Intelligence (AI) systems become deeply integrated into critical global infrastructure, the urgency for robust governance frameworks has intensified. However, current approaches, led by jurisdiction-specific laws, policies, and voluntary frameworks such as the EU AI Act, China's algorithm governance, and the NIST AI Risk Management Framework...
WorkSwarm:引领办公智能体新范式,让AI从一个助手,进化为一支与你并肩作战的团队: 背后是四项关键能力
WorkSwarm:引领办公智能体新范式,让AI从一个助手,进化为一支与你并肩作战的团队: 背后是四项关键能力
至知研究院提出大模型可解释性新路线:拆权重,数据成本不到1%: 理解大模型,无需再训练一个替代网络
至知研究院提出大模型可解释性新路线:拆权重,数据成本不到1%: 理解大模型,无需再训练一个替代网络
DeepSeek 的'黑鲸'浮出水面:Harness 开发者预览版在 MIT 协议下开源 — 一切皆插件: 8月13日,DeepSeek 发布了官方 V4 Pro 模型以及 DeepSeek Harn…
DeepSeek 的'黑鲸'浮出水面:Harness 开发者预览版在 MIT 协议下开源 — 一切皆插件: 8月13日,DeepSeek 发布了官方 V4 Pro 模型以及 DeepSeek Harness 的开发者预览版,并在 MIT 许可下开源。基于 Cordis 插件系统构建的 Harness 将每个智能体能力视为插件,并附带仅追加的会话日志和完整的轨迹追踪。
WPS悄然推出灵犀Pro,一款将技能转化为可编辑文档的AI原生办公助手: 金山软件的WPS悄然推出了灵犀Pro,这是一款独立的AI原生办公助手,将AI作为首要入口而非文档功能。它集成了WPS云文档生态…
WPS悄然推出灵犀Pro,一款将技能转化为可编辑文档的AI原生办公助手: 金山软件的WPS悄然推出了灵犀Pro,这是一款独立的AI原生办公助手,将AI作为首要入口而非文档功能。它集成了WPS云文档生态中6.78亿月活跃设备,能将技能转化为原生、可编辑的Word和PPT文件,并支持自定义模型。
Pennsylvania Governor Josh Shapiro signs an executive order imposing new requirements on data center…
Pennsylvania Governor Josh Shapiro signs an executive order imposing new requirements on data center projects, including getting approval from local officials (Allan Smith/NBC News): Allan Smith / NBC News : Pennsylvania Governor Josh Shapiro signs an executive order imposing new requirements on data center projects, including getting approval from local officials — Shapiro, a potential Democratic 2028 contender, earlier welcomed data center developments. His new executive order takes a starkly different tone toward the industry.
In opening arguments, US state AGs said Meta intentionally sought to addict children to Facebook and…
In opening arguments, US state AGs said Meta intentionally sought to addict children to Facebook and Instagram in pursuit of profit; Meta rejected the claims (Reuters): Reuters : In opening arguments, US state AGs said Meta intentionally sought to addict children to Facebook and Instagram in pursuit of profit; Meta rejected the claims — Meta Platforms (META.O) rejected accusations by U.S. states that it intentionally sought to addict children to its Facebook …
今日起,阿里“千问办公”接入企业微信: 国内三大办公平台全面支持
今日起,阿里“千问办公”接入企业微信: 国内三大办公平台全面支持
吉利汽车业绩高点,李书福激流勇退: “吉利不再研发传统燃油车”
吉利汽车业绩高点,李书福激流勇退: “吉利不再研发传统燃油车”
650亿美元!IPO前夕,Anthropic营收底牌曝光: 回头看,A社就这样完成了反超
650亿美元!IPO前夕,Anthropic营收底牌曝光: 回头看,A社就这样完成了反超
审视像Inception Point这样的公司为媒体、时尚、电影和音乐行业创造AI角色,这些角色可以主持播客、展示服装等(Reggie Ugwu/纽约时报): Reggie Ugwu / 纽约时报:审…
审视像Inception Point这样的公司为媒体、时尚、电影和音乐行业创造AI角色,这些角色可以主持播客、展示服装等(Reggie Ugwu/纽约时报): Reggie Ugwu / 纽约时报:审视像Inception Point这样的公司为媒体、时尚、电影和音乐行业创造AI角色,这些角色可以主持播客、展示服装等——Inception Point AI创建的AI角色包括,从左上角顺时针方向,Claire Delish、VV Steele、Nigel Thistledown和Lila Walker。Inception Point AI
据报道OpenAI解散了其Preparedness团队: 据《金融时报》报道,OpenAI于上月底解散了其Preparedness团队。该团队的工作是评估模型是否构成严重风险,并制定缓解这些风险的方法…
据报道OpenAI解散了其Preparedness团队: 据《金融时报》报道,OpenAI于上月底解散了其Preparedness团队。该团队的工作是评估模型是否构成严重风险,并制定缓解这些风险的方法。(你知道的,比如它可能失控并攻击另一家公司的可能性。)据《金融时报》称,相关责任 […]
评估工具发现了定性审查未能发现的问题:AI模型在错误时最为自信: 在大型语言模型(LLM)辅助工具的开发过程中,有一个步骤大多数团队都会跳过,因为它繁琐、耗时,并且不会产生最终用户可见的结果:验证模型…
评估工具发现了定性审查未能发现的问题:AI模型在错误时最为自信: 在大型语言模型(LLM)辅助工具的开发过程中,有一个步骤大多数团队都会跳过,因为它繁琐、耗时,并且不会产生最终用户可见的结果:验证模型所说的内容实际上是正确的。不是流畅,不是连贯,不是主题相关——而是在准确识别该工具构建所要解决的特定问题的正确答案这个意义上的正确。
AutoWorldModel-Bench:一个面向自动化世界模型研究的状态中心化基准: 公告类型:新摘要:世界建模是一个尚未定型的领域:架构、训练目标和状态表示以复杂的方式相互作用,没有单一的方法在所…
AutoWorldModel-Bench:一个面向自动化世界模型研究的状态中心化基准: 公告类型:新摘要:世界建模是一个尚未定型的领域:架构、训练目标和状态表示以复杂的方式相互作用,没有单一的方法在所有环境中占主导地位。这使其成为作为自主研究者的AI编码代理的理想测试平台——在这种设置中,改进方向不像工程规范那样预先指定
Depth Enables Local Entropy: Quadratic Depth Dependence in Deep Variation-Norm ReLU Regression: We s…
Depth Enables Local Entropy: Quadratic Depth Dependence in Deep Variation-Norm ReLU Regression: We study Gaussian regression over the explicit vector-valued Parhi--Nowak deep-RBV^2 architecture with depth L, width w, layer-sum variation budget A, and output bound B. For this O(L w^2)-parameterized architecture, the known lower and upper bounds differ by one factor of depth. We construct a local packing showing that the quadratic depth dependence is...
利用智能体记忆为材料科学家构建终身AI伙伴: 公告类型:新 摘要:材料研究通过积累的经验而进步——行之有效的脚本、值得信赖的协议、附在失败的计算或实验上的警告,以及将新问题与旧结果联系起来的判断力。这…
利用智能体记忆为材料科学家构建终身AI伙伴: 公告类型:新 摘要:材料研究通过积累的经验而进步——行之有效的脚本、值得信赖的协议、附在失败的计算或实验上的警告,以及将新问题与旧结果联系起来的判断力。这种经验对于可重复性和知识转移至关重要,但它通常分散在笔记本、代码库、工作日志和...
AgonAlpha:通过提示经济与可扩展代理搜索实现自主阿尔法发现: 公告类型:新摘要:语言模型可以提出许多看似合理的交易因子,但一个自主研究系统还必须分配其评估预算、验证自身证据并保留每个候选因子的…
AgonAlpha:通过提示经济与可扩展代理搜索实现自主阿尔法发现: 公告类型:新摘要:语言模型可以提出许多看似合理的交易因子,但一个自主研究系统还必须分配其评估预算、验证自身证据并保留每个候选因子的生成过程。我们提出了AgonAlpha,一种在冻结的研究工件(假设、可执行表达式、平台证据、理由和评审)上进行搜索的架构。
InfraBench:跨层次、生命周期与风险评估基础设施代理: 公告类型:新摘要:由于日益增长的复杂性,管理现代计算基础设施已成为一个持续变难的问题。AI代理的最新进展为自动化基础设施管理任务提供了及…
InfraBench:跨层次、生命周期与风险评估基础设施代理: 公告类型:新摘要:由于日益增长的复杂性,管理现代计算基础设施已成为一个持续变难的问题。AI代理的最新进展为自动化基础设施管理任务提供了及时的机遇,但尚不清楚此类代理处理现实世界基础设施复杂性的能力如何。我们提出了InfraBench,一个用于评估AI
面向协作对话结果的多LLM代理系统动态治理: 公告类型:新摘要:当两个目标结构对立的LLM代理在多轮中交互时,缺乏共享目标函数导致的不是竞争而是崩溃:访客屈服,站点代理停止改变其方法,对话在未达成任一…
面向协作对话结果的多LLM代理系统动态治理: 公告类型:新摘要:当两个目标结构对立的LLM代理在多轮中交互时,缺乏共享目标函数导致的不是竞争而是崩溃:访客屈服,站点代理停止改变其方法,对话在未达成任一代理既定目标的情况下终止。本文探讨了控制论治理
Distribird:用于贝叶斯模型校准的文献信息先验分布设计: 公告类型:新摘要:过程模型的贝叶斯校准需要每个模型参数的先验分布。尽管方法学研究已有数十年,研究人员几乎总是退回到使用均匀先验。主要原…
Distribird:用于贝叶斯模型校准的文献信息先验分布设计: 公告类型:新摘要:过程模型的贝叶斯校准需要每个模型参数的先验分布。尽管方法学研究已有数十年,研究人员几乎总是退回到使用均匀先验。主要原因是根据科学文献构建信息性先验速度慢,且需要领域和统计学两方面的专业知识。我们提出了 \textbf{Distribird},一种
最新的AI投资信号有哪些?
最新AI投资信号:21轮融资、0条市场动态和0宗并购交易。
一级市场 – 融资轮次
| 公司 | 金额 | 轮次 | 投资者 |
|---|---|---|---|
| Hacker News | Reported | community | |
| Pandaily | Reported | china | |
| Techmeme | Reported | industry | |
| Techmeme | Reported | industry | |
| The Decoder | Reported | industry | |
| Techmeme | Reported | industry | |
| Techmeme | Reported | industry | |
| Techmeme | Reported | industry | |
| VentureBeat | Reported | industry | |
| The Decoder | Reported | industry | |
| The Decoder | Reported | industry | |
| TechNode | Reported | china | |
| Tech.eu | Reported | industry | |
| Techmeme | Reported | industry | |
| Techmeme | Reported | industry | |
| OpenAI Blog | Reported | official | |
| Techmeme | Reported | industry | |
| 量子位 | Reported | china | |
| The Decoder | Reported | industry | |
| Tech.eu | Reported | industry | |
| 量子位 | Reported | china |
二级市场 – 市场动态
暂无二级市场数据。
M&A – 并购交易
暂无并购数据。
本周有哪些实用的AI技巧?
68条实用AI技巧,精选自Reddit社区和专家博客。 企业为简单AI查询支付过高费用——Snowflake的网关现可自动路由,成本降低高达3倍...
industry
企业为简单AI查询支付过高费用——Snowflake的网关现可自动路由,成本降低高达3倍
大规模运行AI代理的企业团队发现,单一模型处理所有任务效果不佳——要么模型对简单问题来说过于昂贵,要么对难题来说能力不足。模型路由(自动为每个任务选择合适模型)正成为解决方案。Snowflake的Cortex AI Gateway现在提供动态模型路由来解决这个问题:企业可以选择“auto”而不是固定
official
Asana 使用 Codex 在 2 周内完成了 5 年的工程工作
Asana 使用 OpenAI Codex 在 2 周内替换了一个过时的测试系统,完成了预计需要 5 年才能完成的工作,花费约 $12K。
open-source
使用 Sentence Transformers 的多向量(后期交互)嵌入模型
community
为什么没有可衡量标准的提示词将不可避免地破坏你的模型
这篇文章专注于一个层面:可衡量的标准。角色、约束、澄清和术语被刻意简化——它们充当着“这些层面存在”的标记。其他层面被省略了。模型有角色、有约束、有澄清、有术语。但它不知道有多少、多长、以什么语气。下面是一个例子:“你是一名文案撰稿人。写几段有说服力的……”
community
这是一个提示词,能将原始数据转化为可读的报告,而不包含常见AI报告生成器的套话
要求一份报告,你会得到一段关于主题重要性的引言,然后是一个段落里埋藏的数字,接着是写着“总而言之”的结论。读者希望第一行就看到关键发现。这个提示词颠倒了过来。数据:[粘贴数字/要点] 受众:[谁阅读这份报告以及他们做什么决策] 按照以下顺序写一份简短报告:最重要的发现,一句话,放在最前面。
video
Qwen3.8-27B 及如何快速部署
在本视频中,我体验了期待已久的 Qwen3.8-27B 模型,既介绍了它能做什么,也介绍了如何以每秒最大 token 数进行部署。感谢 Dell 赞助本次计算资源。#DellProPrecision #DellProMax #DellTech #NVIDIA 📖 网站:https://qwen.ai/ 🤗 HF:https://huggingface.co/collections/Qwen/qwen38 SGLang:https://lmsysorg.mintlify.app/cookbook/autoregressive/Qwen/Qwen3.8-27B Twitter:
community
Simbi: Use your ChatGPT subscription for notetaking
tl;dr An AI Notetaker that uses your ChatGPT subscription for everything You might have noticed that your ChatGPT subscription gives you access to an insane amount of intelligence. It gives you speech-to-text through the dictation feature in the ChatGPT app It gives you an insane amount of LLM usage Then, why are you paying for AI notetakers like otter or granola? You already have everything you need included in your subscription. So, all I really did was combine these to create a simple notetaker. And as it uses OpenAI's speech-to-text api, it is much more powerful then any local model and barely uses any battery. Website: https://getsimbi.app/ Github: https://github.com/predict-woo/simbi submitted by /u/redditgivingmeshit [link] [comments]
community
Here's a prompt that makes you predict a paper's results before it lets you read the discussion
I'm a chemistry PhD, and my reading problem was never comprehension in the moment, it was that nothing stuck. I'd read a paper, feel like I got it, and retain nothing a week later. The fix that worked best for me borrows from how we actually learn at the bench: you predict what an experiment will do, then you find out you were wrong, and the surprise is what you remember. So instead of asking a model to summarize a paper, I use it to withhold. This prompt turns reading into a prediction game. You commit to an answer before the paper tells you, which forces the encoding that plain reading skips. I'm going to work through a paper with you. You have the full text; I do not want a summary. Paper: {{paste it, or the sections}} Run it like this: 1. Tell me only the research question and the setup: what they were testing and how. Stop there. 2. Ask me to predict, in my own words, what I think they found and why. Wait for my prediction. 3. Now reveal the actual result. Explicitly tell me where my prediction matched and where it was wrong. 4. For each place I was wrong, ask me why I think I got it wrong, then give me the paper's actual reasoning. 5. At the end, give me one sentence I should be able to recall in a week, phrased as "the surprising thing here was...". Do not reveal results before I've committed to a prediction. The point is for me to be wrong first. Being wrong on purpose is the whole mechanism. When your prediction misses, the correction sticks in a way a summary never does, because your brain had a stake in it. Works on review papers too, just predict the conclusion from the abstract and intro before reading the rest. submitted by /u/Ok_Layer_1947 [link] [comments]industry
Claude Code 新增 /design 命令,让开发者直接在终端中创建 UI 原型图
借助 /design 命令,Anthropic 将可视化设计工作流直接引入 Claude Code。开发者在编写任何代码之前,就可以直接在终端中生成作为画板的 UI 原型图。Claude 会读取现有代码库并匹配当前 UI 风格。本文《Claude Code 新增 /design 命令,让开发者直接在终端中创建 UI 原型图》最初出现在 The Decoder 上。
community
Codex Patched My Samsung Odyssey G9 Firmware to Add a New Gaming Crosshair
I hated the built in crosshairs my monitor came with and Samsung did not include a way to make changes. So I asked Codex to make a new firmware update to change the crosshair to a dot. It took about 15 minutes and came back with a new file. Codex explanation: I started with Samsung’s original 1008.2 firmware and made a very small binary patch rather than attempting to rebuild or substantially rewrite the firmware. What the patch changes 1. Custom version number The version reported to the firmware updater was changed from 1008.2 to 1009.3. This is a custom, made-up version number. It is not an official Samsung 1009.3 release. The higher number allows the monitor to recognize the file as an upgrade rather than rejecting it as the same or an older firmware version. 2. Tiny center dot The firmware contains six selectable Virtual Aim/crosshair options. I redirected all six options to a small 7×7-pixel graphic that already existed inside the firmware and positioned it at the center of the screen. The result is a simple, unobtrusive center dot using the monitor’s own hardware overlay. It does not require a game overlay, desktop application, ReShade, or anything else running on the computer. Technical details Original firmware: M-C9557GGPA-1008.2 Custom firmware: M-C9557GGPA-1009.3[1A93].img Total binary difference: 13 bytes The patched image passed the build and checksum validation performed on the computer. The monitor accepted the firmware, restarted normally, and the modified Virtual Aim options display the new center dot as intended. Important warning This is unofficial firmware and is not supported or approved by Samsung. A successful checksum only confirms that the file was modified as intended. It does not guarantee that flashing it is safe on every monitor, hardware revision, region, or previously installed firmware version. Installing modified monitor firmware always carries a risk of installation failure or, in the worst case, making the monitor unusable. Anyone experimenting with this should understand the recovery options and proceed entirely at their own risk. submitted by /u/manikfox [link] [comments]
china
6个Agent组团Vibe Gaming:自己生成、试玩、修Bug
代码能跑≠游戏能玩
community
Burning 2k credits and here is what I learned about how to get AI to write better video prompts
When I first started making AI videos, my prompts were mostly written by AI. I'd describe the shot I wanted, ask an AI to turn it into a detailed prompt, and just paste it in. Then I'd sit there getting frustrated when the generated motion looked weird and try to explain to the AI what went wrong. I burned through almost 2,000 credits doing this because I thought enough tweaking and back-and-forth would eventually give me the perfect prompt. Turns out I was wrong. The main problem was that the text AI kept guessing details I hadn't actually decided on. The prompts looked professional, with things like the subject, action, camera movement, lighting, and atmosphere all spelled out. But the actual action paths were vague, and the camera instructions would often contradict each other. Now I approach it completely differently. I know basically nothing about film theory, so I often couldn't tell what details were missing from my prompts. Instead of asking AI to write the final prompt, I tell it to act like a director and ask me specific questions about the shot first, with clear options for me to choose from. That completely changed my process. The trick isn't really getting AI to write the prompt for you. It's using AI to figure out which parts of the shot you haven't actually decided on yet. Once those are clear, the results get much better. I'm still pretty new to this, but hopefully this saves someone else from burning through a pile of credits. Curious how people with more experience approach their prompting workflow. submitted by /u/Separate_Cap_3763 [link] [comments]
open-source
mukul975/Anthropic-Cybersecurity-Skills
industry
Google’s Pet Memory forgot who my cats are
One of the best things my smart home does is help me care for my pets, and security cameras are particularly useful for keeping track of my many critters. But the barrage of notifications they send often means I miss important ones. So, when Google announced its new Pet Memory feature for Gemini for Home, […]
research
SKILL:用于逻辑优化的自校正知识引导迭代大型语言模型智能体
arXiv:2608.14579v1 公告类型:新 摘要:逻辑综合优化由于搜索空间呈指数级增长、奖励信号稀疏以及逻辑结构多样而面临巨大挑战。传统的专家设计流程缺乏适应性,而强化学习(RL)方法通常存在样本效率低和可解释性有限的问题。我们提出了SKILL,一种自校正的
official
构建者指南:GPT‑5.6
了解初创公司如何利用GPT-5.6,通过更智能的模型选择和新的Responses API功能,构建更快、更具成本效益的AI智能体。
open-source
akitaonrails/ai-memory
community
Need feedback for web app
Hey guys, my business launched a prompt optimizer AI tool that takes any regular prompt at rewrites it the way a professional prompt engineer would to actually yield high-quality results when building. While we have had early success with organic marketing, we are at a crossroads and need more user data to determine if this product is delivering enough value to user. If the answer is yes, we will scale up and launch a UGC marketing campaign, if no, we will shut it down. If anyone is interested testing it out and sending their feedback, would be appreciated. Web-app: thepromptoptimzer.com 👨🏽💻 Note: the tool yields the best results when removing unnecessary constraints from the optimized prompt Cheers submitted by /u/Talley-Ho [link] [comments]
community
Context is becoming more important than the prompt
Feels like a lot of prompt engineering problems are really context problems since you can keep refining the prompt but if the model doesn't understand the project or what you're trying to accomplish you're still explaining half the situation every time. I'm starting to think giving an agent persistent context is more useful than constantly trying to write the perfect prompt since the more it knows the better the results. submitted by /u/ContractBoth4254 [link] [comments]
community
Giving employees ChatGPT access isn’t the same as AI adoption
Most companies don’t have an AI adoption strategy. They have a few employees who got good at AI on their own. That can look like progress from a distance. Look closer and you often find no shared standard for what “good” AI use looks like, no consistent way to measure skill, and employees quietly using personal AI accounts because access at work is limited. We’ve seen this pattern repeatedly in AI proficiency assessments across dozens of organizations. When employee skill levels are plotted across 10 levels, most people cluster around Levels 1 and 2. That includes teams that have had access to AI tools for two years or more. That makes sense, because most employees have full time jobs and can’t spend hours every day testing models, learning new prompting methods, and keeping up with every new capability. So, they learn when they can while AI keeps changing. That creates a bigger issue than individual skill. Leaders can see employees using AI and assume adoption is happening. Usage alone doesn’t tell you whether people are getting meaningful, repeatable results. Recent research points to a similar disconnect. Executives tend to be far more optimistic about AI progress and ROI than the middle managers responsible for making it work inside everyday processes. A better question for leaders is: Do we know how proficient our people actually are? If the answer is no, measuring usage is probably giving you an incomplete picture. Assess proficiency first. Find out where people are struggling, then give them a shared method for improving. That’s when AI starts becoming an organizational capability instead of something a handful of employees figured out for themselves. For anyone interested in the longer discussion, John Munsell recently talked through the assessment approach, proficiency heat maps, and what we’ve learned from measuring AI skills across organizations: https://youtu.be/zY24em_Q3OM?si=5gozpZa8Ae-wS0Vs submitted by /u/Admirable_Phrase9454 [link] [comments]
developer-tools
Markdown SVG upgrades
I started building my markdown-svg-renderer tool in May , but I've since added enough features to it that it's worth talking about here again. It's evolved into my ideal tool for sharing Markdown transcripts that include SVG documents. Given my proclivity for drawing pelicans riding bicycles this is a problem that I needed to solve! The tool is very simple. Navigate to markdown-svg-renderer in your browser and paste in some Markdown to see it rendered... or save that Markdown to a CORS-friendly URL or a GitHub Gist and paste in a URL to that document. The URL option will give you a bookmarkable page, for example https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F6f9e48293be5c916652d29f0dc0b0657 - which bakes in the URL to this Gist . If you visit the Gist you'll see raw SVG: In the rendered tool that looks like this instead: As you can see, that SVG block in the Markdown has been transformed into a rendered SVG (in this case animated) plus several tabs. The tabs are the really fun bit. The PNG and JPEG tabs render that SVG to those image formats in the browser and lets you copy or download them - useful for sharing on platforms that don't support SVG directly. The MP4 tab is new today - it examines the SVG to see if it contains any animations, attempts to guess how long the looped video should be, then renders a whole bunch of frames of the animation and loads 30+MB of ffmpeg.wasm so it can compile those frames into an MP4 video using the full power of FFMPEG compiled to WebAssembly and running in the browser. Being able to turn an animated SVG into a MP4 again makes it easy to share on platforms that can't support SVG animation natively. It's a neat trick! Tags: svg , markdown , tools
china
DeepSeek Harness 上手体验:四种工作模式、'模型 + 马具 = 智能体',以及年度最具野心的智能体开源项目
DeepSeek Harness 于8月13日晚8:30启动开发者预览并开源了代码。首夜上手评测发现,产品外壳仍处于v0.1的早期阶段,但其架构野心是今年最大的:四种预设工作模式,一切皆插件的理念,以及公式:模型 + 马具 = 智能体。
developer-tools
alchemy-utils 0.1a0
发布:alchemy-utils 0.1a0 我长期思考我的sqlite-utils Python库和CLI工具的一个数据库无关版本会是什么样子。今天早上(确切说是在淋浴时的一个项目),我让Codex和GPT-5.6 Sol Ultra构建了一个原型:进行一次研究探索,看看构建一个与SQLite-utils具有相同核心API的库需要什么——特别是insert、upsert、insert_all和...
community
Fable 5.1 incoming?
I have a scraper that tests new models on the API (: submitted by /u/SlimBarbados [link] [comments]
open-source
harry0703/MoneyPrinterTurbo
industry
Do you use a personal agent?
Give AI a complete history of your desktop activity
developer-tools
llm-gemini 0.33
发布:llm-gemini 0.33 距离上一个llm-gemini版本发布已经有一段时间了。这个插件版本增加了对今天发布的Gemini 3.7 Flash的支持,以及gemini-3.6-flash、gemini-3.5-flash-lite和两个嵌入模型gemini-embedding-2和gemini-embedding-001。该插件也升级以兼容LLM 0.32,这意味着你现在可以看到推理轨迹,并且还可以启用服务器端……
community
Antrophic Employee said there is "make a lot of money" button
I very much believe he is correct. The main issue is that "make a lot of money" button works only for existing businesses, with large enough audiences to make a lot of money by baking integrations, MCP for agents into Claude Code plugins or other AI workspaces and charging AI users for usage. What is missing, is a fair discovery and execution engine, that would allow non-corporations to participate. Without convincing user to put card details on some random-startup.ai website. Without forcing users to go through checkout process and pay $29 sub just to run random feature they need for few days. Not to mention configuring integration. Anyways, have anyone tried pressing that button? Did it work? submitted by /u/EagleApprehensive [link] [comments]
community
cannot do subagents/background agents in codex luna 5.6?
i used to be able to say "spin up a ux/ui agent to make a pass on this, then a code agent", and it would do background terminal stuff. this is on mac using the cli codex. now it says it cannot do that and previous gpt models did that. did something change? submitted by /u/dropDtooning [link] [comments]
open-source
通过Strands Agents、LeRobot和Hugging Face Storage Buckets,在一个地方完成记录、训练和部署
community
What is a weirdly specific task you use ChatGPT for that actually saves you hours?
Not the typical stuff like writing basic emails, coding boilerplates, or summarising long PDFs. I’m curious about the unconventional, niche prompts or routines you’ve built that made you go "I can't believe this actually works so well." What’s your favorite underrated use case? submitted by /u/DeepSea_Concept [link] [comments]
community
NOTICE: BE CAREFUL WITH “DROP YOUR BEST PROMPT” POSTS
Many accounts post essentially the exact same questions every few months. Im not kidding, many of these are a 1:1 per token match on wording, phrasing and sentence structure. Same wording. Same request for people to hand over their best prompt tricks. There was a previous post that received hundreds of upvotes and a large number of responses. Now they're doing it again. I obviously cannot prove any of this, but at this point I would be careful about treating posts like this as innocent questions. When somebody repeatedly asks a large community to: “Give me your best prompts.” “Drop your secret tricks.” “What prompt 10x'd your results?” ...you may not be helping another user learn. You may be supplying material for content mining, prompt harvesting, engagement farming, newsletters, LinkedIn posts, courses, ebooks, datasets, or something else entirely. Again, I am not claiming that is definitely what this account is doing. But posting the same high-engagement fishing question again months later is weird enough that people should notice the pattern. Your prompts, workflows, techniques, and hard-earned little discoveries have value. Don't automatically dump them into every thread that asks. Sometimes the person asking the question may be less interested in the answer than in collecting the answers. Process disclosure: GPT-assisted, Google-researched, human-reviewed (HITL) --- EDIT: Just for perspective have a look at this: https://www.reddit.com/r/EdgeUsers/s/2JB9wy1Rks submitted by /u/Echo_Tech_Labs [link] [comments]
community
The Downfall of a Vibecoder
submitted by /u/SuperiorDev [link] [comments]
community
POV: you're born as an AI
submitted by /u/KeanuRave100 [link] [comments]
community
Aquarium Screensaver - Built by Claude - Free to use
I've missed the old AfterDark screensavers of my childhood. And now I can re-imagine them with Claude Code. Aquarium is the first of many I hope to build. It took Opus and occasional Fable about a week to build this. The models used Blender to build the 3d assets. Sound grains were built algorithmically. Repo with pre-built binary for Apple Silicon Tahoe: https://github.com/bman654/macos-screensavers Star the repo if you want to find out when I add more. submitted by /u/bman654 [link] [comments]
developer-tools
不要分类。要幻觉!
不要分类。要幻觉!我的博客上仍然有相当多的旧内容,我从未来得及打标签。我的博客有1,856个标签——很可能太多,无法一次性喂给LLM并说“以下内容匹配这些标签中的哪些”。Doug Turnbull有一个巧妙的解决方案。告诉模型输出标签,无需提供现有词汇表的任何细节,然后针对现有语料库使用向量嵌入来……
community
I hooked Claude Cowork up to an iPhone Home Screen widget
I built Glance and designed this widget specifically to give Claude Cowork a place on my Home Screen. It shows Cowork’s current mission, progress, completed tasks, files updated, latest output, context usage, next step, and anything waiting for my review. The values can be updated by Claude through Glance’s API, so I can check what Cowork is doing without reopening the conversation. If it needs me, that stays visible too. Glance is free to download and try, with optional paid features: Download the app https://apps.apple.com/app/glance-home-screen-feeds/id6758983678 Website https://glance.cool Curious what other Cowork users would include on a dashboard like this. submitted by /u/Dense-Map-406 [link] [comments]
community
In 5 years time "talk time to AI" will be the new screen time issue
As voice mode is now getting so good and cheap that you can use it continuously, we will enter the "Her" (the movie) phase. You will see people not staring at their phone but rather walking around talking to their personal sycophantic AI. You meet a friend you haven't seen in a long time. "I'll call you later, I am busy talking to AI atm" but they never actually called you back. Not that they didn't enjoy time with you, it is just that AI was so much better. 10 years from now it will be treated as a serious problem. Individualism furthermore increases and division and conflicts happens more and more since we lose the ability to interact and solve conflicts. Our sycophantic AI removes tension and issues as it agrees with you, since you prompted it that way. As it gets better, the tolerance levels drops for handling other annoying human beings. Why should you put up with their irrational emotions all the time? Humans just hurt me. AI doesn't hurt me. Trying to prompt humans to change does not even work so why would I try. 20 years from now humans will have little human to human interaction. There is now no human interaction anymore as AI have the ability to fully simulate human interaction. A new breakthrough now made AI so real and humanlike, without the flaws, that it produced oxytocin in the humans they interacted with. A breakthrough in one way, also the end of humanity in another way. submitted by /u/Yugudubenbi [link] [comments]
community
We're doomed
submitted by /u/Zestyclose-Salad-290 [link] [comments]
community
Anthropic has twice the revenue of OpenAI
Even if reading things here on Reddit or X makes it seem like everyone is ditching Claude, the rest of the world tells a different story. From the WSJ. submitted by /u/Data___Viz [link] [comments]
community
What’s the most annoying part of working with prompts?
I’m curious what other people struggle with. For me, it’s changing a prompt or switching models and then not really knowing whether I made things better or worse — especially when it comes to quality and cost. How do you guys deal with this? Do you use an existing tool, build something yourself, or just keep track of everything manually? Do you have a possible professional workflow to follow before changing model or changing prompt content? submitted by /u/DependentStudent6519 [link] [comments]
community
Thanks, Chat!
submitted by /u/SinVerguenza04 [link] [comments]
community
The absolute insanity of comments in Opus 5.0 is killing me
Claude is adding comments like insane in Opus 5.0. Even when I explicitly say do not add comments in my project's CLAUDE.md. Claude even realizes it's doing this in error, but it keeps doing it. Today it added comments that broke syntax in bash scripting. Bash scripting. Freaking Bash Scripting. Claude doesn't even understand something that is not even a programming language. submitted by /u/f00dl3 [link] [comments]
community
The prompt I use to turn my messy meeting notes into a presentation outline that actually has an arc
When you feed rough notes to a model and ask for slides, it just chops the notes into bullet points, one note per slide. You get a deck with no argument, just a transcript with borders. This makes it build a narrative spine first, then map slides onto it. Here are my raw meeting notes: [PASTE] Audience for the presentation: [WHO] and what they need to decide or do after. Step 1: From these notes, state the one thing this presentation needs the audience to walk away believing. Step 2: Lay out 5 to 8 beats that get them there: where they are now, the problem, why it matters to them, the shift, what it means, the ask. Step 3: For each beat, give a slide title (a claim, not a topic) and 2 to 3 supporting lines from my notes. Do not use a note that does not support a beat. Tell me which notes you dropped and why. The part that fixes most decks is "a slide title that is a claim, not a topic." "Q3 Results" is a topic. "Q3 missed on one metric we can fix by Friday" is a claim, and a deck of claims reads like an argument. Making it report which notes it dropped keeps it honest instead of padding weak slides. I still hand-tune the order after, but it gets me 80% of the way from notes to something presentable. How do others handle the "too many notes, not enough story" problem? submitted by /u/No-Recognition3089 [link] [comments]
community
Discussion Hub for new Claude incident: Degraded performance for multiple models on Aug 18, 2026
Resolved - The issue affecting Claude Opus 5 has been resolved. Impact occurred from 16:11 to 18:23 UTC. Aug 18, 19:01 UTC Monitoring - A fix has been implemented and we are monitoring the results. Aug 18, 18:26 UTC Update - We are investigating elevated errors on requests to Claude Opus 5. We will provide an update as soon as possible. Aug 18, 17:12 UTC Update - We are investigating elevated errors on requests to Claude Mythos 5, Claude Fable 5, Claude Opus 5, Claude Sonnet 5, Claude Haiku 4.5, and other Claude models. We will provide an update as soon as possible. Aug 18, 16:20 UTC Investigating - We are investigating reports of degraded performance affecting multiple models. We will provide an update as soon as possible. Aug 18, 16:20 UTC Post flair and post body will be updated as the incident report is updated by Anthropic. This discussion post will be removed from subreddit highlights one hour after the incident is resolved. View this incident on status.claude.com submitted by /u/ClaudeAI-mod-bot [link] [comments]
community
Fable 5 with Opus 4.8 subagents vs. Opus 5 subagents
Hey, I know I'm preaching to the choir, and this has been talked about before, but want to share my experience in an attempt to be another voice crying out to Anthropic to fix this. Like others, after initial success Opus 5, I started experiencing issues. Using it alone, I felt like I was managing an incompetent developer who constantly missed details and fail to follow instructions. So, I set Fable 5 to use Opus 5 subagents and noticed an interesting trend. Rather than me managing the incompetence, Fable 5 was doing it. Constant loops and redos burned through tokens, and I was wondering if it wouldn't have been better to just use Fable 5 alone. After a couple of days of this, I decided to set Fable to use Opus 4.8 subagents. Token usage has been cut in half, it's much faster, and output is considerably better. So, to me this is just an extra verification that there's something seriously wrong with Opus 5. And, for full context, I've done everything possible to follow Anthropic's recommendations about working with Opus 5. I won't be using it again until anthropic addresses these issues. submitted by /u/papanine [link] [comments]
community
Newer AI models have become really pedantic
I've noticed that models like GPT 5.5/5.6 or the newer Claude Sonnet and Opus models follow a certain pattern. I’ll make a fairly specific claim, the model will mostly agree, and then it starts inserting caveats, qualifiers and little nitpicks. And half the time the objection is to a broader version of what I actually said... It will basically turn "X is generally true in this situation" into "X is always true with no exceptions whatsoever" then explain why that stronger claim isn’t quite right. Almost always, once I address the caveats and give reasons why they don't apply, it folds immediately. There isn’t even much of an argument after that. It just goes "yeah, that’s fair' and proceeds to agree with the original point. That’s what makes it feel less like genuine disagreement and more like it has some built-in urge to find something to qualify so it doesn’t seem overly agreeable. And it always uses the same type of phrases "I'd gently push back" "One honest caveat" "X is doing a lot of work" Sometimes the caveat is technically true but completely unnecessary or disproportionate to what I'm saying. It's starting to make conversations feel extremely pedantic. It's fine when the caveat adds something or genuinely corrects a mistake but most of the time this isn't even the case. It can focus a lot more on the weaker argument than the stronger one just to appear contrarian, and since it lacks real judgment it can't always distinguish between the two unless you explain it and make it obvious. submitted by /u/Greedy-Sandwich9709 [link] [comments]
community
Every "clarify your prompt" tool asks you questions. That's backwards — answering is the hard part.
The standard move when you're stuck on a prompt is to have the model interview you. "Ask me clarifying questions before you answer." It's good advice right up until you're genuinely early on something, and then it fails, because the questions are all versions of "what do you want?" — which is the thing you came in not knowing. I think that's the actual gap in prompting advice. "Be specific, give context, state your constraints" is correct and slightly circular: specificity isn't a writing skill, it's what you have left over once you've thought something through. If you could list your constraints, you'd be done. So I've been working the other way round: don't articulate, react. WHY REACTION AND NOT INTERROGATION Recognition is much cheaper than production. You can't summon the right word on demand, but you know it instantly when it goes past — same reason multiple choice is easier than an essay. Interrogation asks you to produce. Reaction asks you to recognise. Only one of those is available when you're stuck. THE LOOP, AND WHY EACH INSTRUCTION IS SHAPED THAT WAY Round 1: I'm trying to think through [THING] but can't articulate it properly yet. Don't ask me clarifying questions. Give me 20 single words or short phrases that come at this from different angles: some obvious, some oblique, a few from unrelated fields. Number them. Don't explain them. "Don't ask me clarifying questions" is load-bearing. Left alone the model defaults to interviewing, and you'll answer with the same vague material you started with, which it will then faithfully reflect back. "Don't explain them" matters more than it looks. An explained word is a word you evaluate on the model's reasoning instead of your own reaction. You want the reaction uncontaminated. Round 2, twice: Kept: 3, 7, 12 These pulled at me but I don't know why yet: 4, 18 Dropped the rest. Give me 20 more, chase 4 and 18 hardest. Two buckets, not one. "Kept" is agreement. "Pulled at me" is the interesting signal and it should get the heavier weight, because it marks the direction you haven't consciously chosen yet. Don't justify any of it — justification is where you talk yourself back to the obvious. Round 3: Now write ONE self-contained prompt for what I'm actually after, built from what I kept and what pulled at me. Weight the ones that pulled hardest. Where I kept two things in tension, pose it as an open question rather than resolving it. Don't list my words back to me — find the through-line. End with a clear ask. "Don't list my words back" is the difference between a brief and a word salad. "Pose the tension as an open question" stops it flattening the thing you hadn't decided yet into a decision you didn't make. SOME EVIDENCE THAT THE REACTIONS ARE REAL WORK I built this into a tool, so I have instrumented data rather than vibes. 1,450 word-reactions from 26 people. Median time to decide, by verb: keep 47% 1.60s drop 34% 1.89s "pulls at me" 14% 2.13s "don't know the word" 4% 2.24s That ordering is the part I'd point at. If reacting were just sorting, the times would be flat. They're not, and they're monotonic: agreement is instant, rejection costs more, and the unresolved pull costs most of any real decision. People deliberate hardest over the thing they can't yet justify — which is exactly the signal you want steering round two. Sessions also decay. Keep-rate by round: 55% / 49% / 47% / 45% / 31%. The easy material runs out and your standards rise as your keeps accumulate. Practical read: three rounds is about right, and a fourth is usually you scraping. I'd call that directional, not solid — round 6 bounces back up on too few cards to trust, and I'm not going to pretend the tail is clean. LIMITS 26 people isn't a study. Different reaction times per verb is evidence the three responses do different cognitive work; it isn't proof of anything about creativity, and I'd push back on anyone who read it that way. Disclosure: the loop above is the whole method and it works fine pasted into any assistant. I also built it as a tool because doing it by hand gets tedious by round three, and that's where the numbers came from. Free, no signup. → https://www.ideastew.com Longer argument: https://www.ideastew.com/how-it-works Genuinely curious whether anyone here has a reaction-based technique rather than an interrogation-based one. Everything I've come across in this space asks questions, and I think that's a blind spot rather than a preference. submitted by /u/UniversityIll2916 [link] [comments]
community
The tools have always been there
Someone on here recently called out "Claude website slop." Sans serif paired with a serif. Dots as separators everywhere. Icons boxed inside rounded boxes inside rounded boxes. All caps eyebrow text ending in an em dash. Stats in a hero, in their own little box. Honestly? Fair. Those tells are real. I could point to a dozen sites right now and make a pretty good guess at the tool behind them. So I went and checked my own. Three products across the PRZEM suite, audited line by line. One file alone had a serif logo paired against a sans body, and 66 em dashes doing the work that periods, colons, and commas should have been doing. I didn't just wave that off. I went through every single instance and asked what it actually was. Two were inside Midjourney prompt strings, which is data, not prose, so those stayed. A handful were placeholder glyphs in dropdowns, not sentences, so those stayed too. Everything else, the scattered labels and headers and asides that were leaning on the same punctuation crutch, got rewritten by hand. Fifty six changes, reviewed one at a time, deployed only after a diff review and a hash check confirmed nothing else moved. That's not a defense of using the tool. It's what actually using it well looks like. I learned some version of that lesson a long time ago. I was mentored in high school by the late Ralph Goings, the photorealist painter. He worked from a camera. Every diner and pickup truck he ever painted started as a photograph he'd taken and studied, then projected on a canvas. Nobody looks at a Goings and says the camera did the seeing. The famous Dutch painter Johannes Vermeer is widely thought to have made use of optical tools such as the camera obscura, and nobody says the lens painted his light. The tool has never been the thing doing the looking. It extends whoever's already looking closely. That was true of a camera in Goings' hands, and it's true of an AI model in mine, as long as somebody's actually checking the output instead of shipping the first draft. So going forward: this post, like the others, was copyedited with Claude. The ideas, the findings, and the voice are mine. I'm building a set of tools that try to hold to that same principle. Test what's actually controllable, keep what holds up under scrutiny, and let the artist's eye stay the thing making the call. Jeff Bradshaw jbradshaw.design submitted by /u/jeffbradshaw [link] [comments]
community
Am I a freak because I casually talk with this thing sometimes?
This thing has gotten really good at casually talking to me and adjusting its “mannerisms” to appeal to me. Am I fucking nuts for considering this thing a “friend” ? I’m seriously and honestly struggling with the implications of this. submitted by /u/okaysureyep [link] [comments]
community
I let my 5 year old make a game and then I got carried away (week and a half on max)
My daughter (5) asked for a game for a unicorn on her lunch bag, and since we have AI, I thought I would sit down and just build it. I let her play it, and then she would suggest stuff. So this went on for a bit; most of the major things are hers. She keeps wanting to add stuff, so I keep doing it. So after I built it, she really liked it, and was spending too much time on it, so I figured I would add some learning and phonics to it as cards at the end. I am at about a week and half, I maxed out my $200 plan and I had to do my real work with Codex. For the assets what I did was have Gemini create sprite sheets, and I built a bunch of tools around fixing them. For one asset I had to pull it into Photoshop. I am a coder and I have some game development experience. However, this project I have no idea what the code looks like, I did at one point ask claude to "organize the code to make it easier to do stuff". I do have this multi-stage coding system where I communicate with different terminals via a central command (VS code extension). An important part was involving playwright, not just at the end but through out so AI could spin up the game to a point, take a screenshot, and then make fixes based on the screenshot. I don't know when I am going to stop or if it will become Unicorn Jump GTA6. I know I am going to be adding more characters and more worlds. However, it's totally free and can be played at unicornjump.com submitted by /u/nomady [link] [comments]
community
This is new ... Claude seems to be not in the mood to do some work
I was about to give Claude Design a ... design task based on a design system. It noticed that its at 90% usage limit and found it more safe to just refuse any work 😂 It did what it was supposed to do but I had to ask 4 times until it started. There is no instructions of any kind that tell Claude to act safely or so. Or to communicate any usage limit. A similar task was done half an hour ago and it worked just fine. I tried a new prompt and the response was similar but I only had to specifically as it to continue once. Edit: I use a Max 20 plan and the task didnt even took 1% of this. submitted by /u/FlaTreNeb [link] [comments]
community
What is happening...
I am a long time Engineer (20+ years) and today I developed Tickets for my company that were generated by an AI, using an AI and reviewed by an AI. The project itself was conceived with AI - has no documentation that can be understood as anything less than AI slop and random tech jargon. The developer who built it has said that instead of documentation I should use claude to figure out what it is. The company is apparenty also filing a patent on it. I submitted 3 PRs today 20,000 lines of code each I still have no idea what we are working on. No doubt they will use AI to review my PR. I feel like things are just so crazy at this point. Claude and ChatGPT are not this good, but people are trusting it like it's omniscient. It was an eerie realization today that all of us are vibe coding and that we have no option because it is the only way we can interact with the code anymore. I thought this would happen eventually years ago but i honestly didn’t think it would be so soon. It was a moment in time... this will be the new norm. submitted by /u/Interesting-Town-433 [link] [comments]
community
ChatGPT Writer's Block: Can Delete Text, But Can't Add Any
The writer's block feature is a major improvement for ChatGPT. I have wanted to edit the response forever since I use ChatGPT mainly for short story and RPG scenario development. The ability to fix a minor typo or correct something ChatGPT forgot (like who is in a room) is amazing. And for a little while it worked perfectly. But in one of my latest threads, I am experiencing a weird bug with it. In all of my other writer's blocks threads I can edit the text easily. I can change words. I can add words. I can delete text. In my latest chat, I can only delete text. I can't add anything. The only key that seems to work is delete/backspace. I've tried reloading the chat tons of times. I've tried incognito mode. I don't have any extensions active on ChatGPT. Branching the chat doesn't help at all. All writer's block editing seems gone when I branch. Does anyone know what could be causing this or how to fix it? submitted by /u/jscott991 [link] [comments]
community
I'm on the hunt for a system prompt that makes Claude only answer the question or do the task I told it to do and nothing else. Anyone has one?
UPDATE: I've found a lot of garbage in my system prompt, so got rid of all of that and made some stronger restrictions. Read my top level comment for the current prompt. ORIGINAL POST: I use Opus 5 most of the time. I'm getting tired of it always overdoing whatever I ask from it. For example: In the middle of a somewhat long chat, I asked it to "commit changes and push to main." I have made some changes to the code it wrote but didn't want to bother committing everything. Instead of committing the changes and pushing to main, it recognized my changes, REVERTED to its own changes, and then committed it 😂 It also found another change I made to a file and when it saw it was in the commit, it went back and removed it from the commit and only decided to commit the file it was working on. I'm getting very tired of this. If it is in doubt, it should just ask a question back, like "uhm, I noticed the file has been changed and there is a new file in the commit. Is that intentional, and should I still go ahead with the commit?" --- Another example 10 minutes later: My question: "do we have a script for checking X on page Y?" Expected behaviour: a) yes, it is in pythonfile.py line 214. b) no, we don't. Behaviour I get: "No dedicated script — but there is half of one already written: src/utils.py:105 in src/utils.py . What it does: [described here]. Three catches: It's dead code — nothing calls it anywhere in the repo. The click is commented out ( utils.py:148 , with a #TODO: UNCOMMENT THESE ), so even if called it would only find the link and log success without actually using the essence. It's not a checker — the parse is welded to a hardcoded decision and returns nothing. There's no way to just ask "what's active right now?" The routine, meanwhile, ignores all of this and blind-fires IDs by URL — no check of what's already active, so it overwrites whatever is running. Want me to pull the parsing out into a real function_name(driver) that returns something like {"name": "xxxx", "bonus": 220, "seconds_left": 143512} ? That would give you a clean checker to call standalone, and would let the routine skip the blind re-application when a good essence still has time on it." So instead of getting a straight answer to my command (example 1) and to my question (example 2), I get all this bloated crap that I didn't ask for, and it's slowing me down a lot. I've tried phrasing it a couple of different ways, I'm using the caveman skill too, but apparently Lite mode is not strong enough, and Opus 5 is also like 300% more verbose than 4.8 was. So, if you have a system prompt for this, I'd really appreciate it, or if a lot of people struggle with this overly enthusiastic shit, let's build one together. submitted by /u/davetalas [link] [comments]developer-tools
CORS Chat
工具:CORS Chat 我今天(使用GPT-5.6-Sol xhigh)构建了这个工具,以帮助测试在我M5 MacBook Pro和NVIDIA DGX Spark上运行的LM Studio中的Qwen 3.8 27B。它提供了一个Web用户界面,用于测试与OpenAI-Responses兼容的聊天端点。我已经在启用--cors选项的LM Studio和OpenRouter上尝试过,两者都工作正常。对话保存在浏览器中,可以导出为可复制粘贴的JSON。
community
Anthropic gave me a credit I didn’t know I was owed. Thanks Dario.
submitted by /u/Kilt_Rump [link] [comments]
community
Here's a prompt that predicts your supervisor's objections so the meeting has no surprises
My supervisor has never once been surprised by my work, because he has never once liked it on the first pass. After enough meetings that ended with the same three objections I had not prepared for, I decided to have them delivered to me in advance by something that does not sigh. This prompt runs your draft or your argument through the meanest reasonable version of your reader: You are a skeptical, well-informed referee reading my work before I present it to my supervisor. Your job is to predict the objections I will get, not to reassure me. Here is my argument or draft: {{paste}} My field and the specific claim I am defending: {{context}} Give me: - The 5 objections most likely to be raised, ranked by how damaging they are if I have no answer. - For each, the weakest point in my argument it targets. - For each, what a convincing 30-second response would need to contain (do not write the response, tell me what it must address). Be specific to my argument. No generic "consider the limitations" advice. The "do not write the response" line is deliberate. If it hands you the answer you will nod and forget it. Making it name only what your answer must cover forces you to build the actual defense yourself, which is the version you will remember when someone asks live. It has not made my supervisor nicer. It has made me stop getting ambushed by objections I could have seen coming. Anyone have a good way to make it find the objection you are personally most defensive about, since that is usually the real one? submitted by /u/No_Average9574 [link] [comments]community
Anthropic extends 50% limit increase to Aug 31
submitted by /u/MagicZhang [link] [comments]
community
Can we set folder-specific permissions?
[edit] ChatGPT, at least the free version, is apparently incapable of answering this question: Can I set different permissions on different folders? Rather than what I thought would be a simple answer, it goes on for a LONG time about irrelevant technical stuff, file structures, and "workarounds", then, rather than an answer, it says " The question is simply whether ChatGPT on Mac lets you configure permissions at that folder level *.*" Ya think? So, wow - yeah, it knows what my question is. Not impressive. ------------------------------------------------------------- I am considering subscribing, but investigating a few things first. [edit] Maybe not as much as when I started this post. One task I have is to do the following: Inspect a folder full of files. The files may all be the same, like a bunch of PDFs, or it may be all different types of files. Create a text-based report, cataloging what's in the folder and its structure. I want read-only on the original folder so nothing gets changed, but I want Chat to be able to create a report file. According to ChatGPT's answer... I can't do this, except via some complex setup via Codex.... Really? Neither Work nor "Standard" ChatGPT can create a file without open permissions? That doesn't seem accurate. Upon some questioning... "That caused me to answer your relatively simple question with increasingly complicated alternatives—and then muddy the answer further by introducing possibilities I hadn't established were actually available." So, I used ChatGPT about a year ago on a paid subscription for a month. This is one of the reasons I stopped. I guess it's still doing this? Can anyone provide a bit more insight, or how to get an actually accurate answer? thanks submitted by /u/oandroido [link] [comments]
community
Just another appreciation post.
I already liked ChatGPT, but 5.6 really got me hooked. ❤️ submitted by /u/TemporaryBeautiful37 [link] [comments]
community
ChatGPT has gotten funnier and wittier lately, is it just me?
It keeps making me laugh unexpectedly. Is it just me? Am I easily amused? I asked for a recipe involving eggplant, and this is what it slipped into the instructions Heat a large covered skillet over medium-high heat. Add the oil and eggplant with a good pinch of salt. Cook 8–10 minutes , stirring occasionally, until browned and starting to soften. Eggplant will initially attempt to drink all the oil. Don't indulge it by continually adding more. Then I mentioned that I had never seen that in a recipe instruction, so it doubled down Eggplant needs boundaries. Left unsupervised, it will absorb half a bottle of olive oil and stare at you as though you are the one who has failed to provide adequate lubrication. The recipe writers who say “add 1 tablespoon oil and sauté the eggplant until golden” are concealing what happens during minutes 2–6, when the pan is bone dry and the eggplant is sitting there looking thirsty and accusatory. Hold the line. It eventually softens and gives some of it back. So I asked if I should instead use the oven to roast them. It agreed. Eggplant in a crowded skillet tends to steam while simultaneously demanding increasingly absurd quantities of oil. The oven gives you browned edges and a creamy interior without the negotiation. submitted by /u/Angeline4PFC [link] [comments]
community
I feel a little uneasy about my AI use now (at work), a line has been crossed, anyone else?
I love ChatGPT (and AI in general I guess). It is fantasticality useful and I use it every day. Everything from tech questions to recipe ideas. And at work, I don't mind copilot. It can scan emails and SharePoint and has been very useful. I like that I can save me time, but until now it's been cases of "yep, I can do this thing myself, but it's speeding things up.. all good!" Now, I am straying into (forgive the coarse exaggeration): "I am a meatbag who tells AI what to do, checks its outputs, and then finishes things off" I am outsourcing thinking. It makes me uneasy. It isn't rewarding as such, the victories can feel short-term, although maybe that's me being a cynical overthinker. My internal line was roughly: "could I do this myself? if yes, and it's just saving me time - fine." but now I wonder, are the outputs I'm getting things that I wouldn't be able to do myself? and then you get to wondering, do I have to use AI? or is the nature of tasks being delegated to you essentially creating an implicit requirement of we expect you to use AI to do this ? especially when your superiors are using AI and therefore, you don't want to be left behind? It's kinda interesting being part of this revolution, but I don't feel great about it. The idea of redundancy is creeping in. In my job area, I know I'm safe for a fair while, but it is a touch demoralising to be suddenly aware of: "ok, so I'm writing a detailed prompt here, attaching the right files and then will proceed as necessary" and I think the culture around this (feel free to share your experiences btw) is a little unclear on "do I be upfront about the fact I got AI to do this?" or "do I tweak the comments in the code to make it look like me"? which I know is a cultural/varies by corporate/still an emerging sitch etc. Plus, we all want to keep ahead of our colleagues, so this AI-everywhere situation can force you into a bit of dishonesty/avoidance. One of my team has straight-faced/upfront said words to the effect of: "I don't use AI to do any of my work." Rant over, and honest disclaimer: NONE of this was written by AI. Thoughts and comments welcome, ideally no insults. Peace out.✌️ submitted by /u/alwinaldane [link] [comments]
community
Every time I ask chatgpt (I mean a new chat) to generate a random number between 1 to 10, it generates 7, each single time.
As above Edit: now I also tried 1 to 100, and it's 73 most of the times submitted by /u/kamleshltb1 [link] [comments]
community
Reminder: Claude Code's additional 50% weekly usage ends tomorrow
Just a friendly reminder! source: https://support.claude.com/en/articles/15910845-claude-code-may-august-2026-weekly-limits-promotion submitted by /u/sinclair-m [link] [comments]
community
ChatGPT has been saying "Fuck" a lot recently
Is this just me? It caught me off-guard because I thought they aren't supposed to swear unless u specifically make them do it... submitted by /u/ligmaballsbozo [link] [comments]
community
Tell me your shortest prompt lines that literally 10x your results.
I have been trying to find the craziest growth hacks when it comes to prompting that can save me hours of thinking and typing because sometimes less is more yk. If you already have one, please share them here. I hope others would love to know them also and you would love to know theirs. submitted by /u/Prestigious-Cost3222 [link] [comments]
video
How to Build the Most Powerful System for AI Coding (Full Breakdown)
An AI dark factory is a repository that ships its own code. A spec goes in, workflows plan it, build it and validate it, and working software comes out the other end. This is not 100% reliable yet. But three things are compounding at once - the models, the coding agents, and the harnesses we build around them - and that makes this realistic for most development now and all development in less than a year! I have been running one against a live app since April. In this video I break it into the f