Insurers claim AI is already increasing healthcare costs
Blue Cross Blue Shield says hospital use of AI tools led to an additional $942M in healthcare spending over a two-year period.
Señal sobre ruido
Una vista priorizada de IA, software, investigación y herramientas desde fuentes primarias y comunidades técnicas.
Actualizado
Señal principal
Priorizado de las últimas 36 horas
Blue Cross Blue Shield says hospital use of AI tools led to an additional $942M in healthcare spending over a two-year period.
Hello HN, It started as an experiment: can Claude play chess properly if it uses vision instead of PGN notation? Somehow it can. The next experiment was to see whether Claude + Stockfish could explain a game. Somehow it can too. A few sessions later, I had a system that takes my live audio notes (or text, for that matter) and a vague instruction like "analyze my last lichess game", and gives me a commented video of the game. The result is not perfect and it takes time to deliver (an hour or so), but for me it is a much more pleasant and memorable experience than clicking around Stockfish branches. It burns tokens, so make sure you have enough…
After obtaining an interactive avatar and training it to discuss venture fraud, I have mixed feelings about making AI clones of ourselves.
When AI leaders at OpenAI and Anthropic started talking about “pacing the frontier,” maybe someone should have asked: what pace? Now it’s turned into model drop week for both companies as Anthropic rolled out Opus 5.5, followed by OpenAI’s GPT-6 model updates just 90 minutes later. But the company that stole the spotlight was Meta, whose personal AI agent Muse is reportedly outpacing ChatGPT’s early numbers and is headed for smart glasses and […]
GitHub Copilot in Slack and Microsoft Teams now gives you more context, more control, and a clearer path from conversation to GitHub work. Whether you’re sharing files in Slack or… The post Updates to GitHub Copilot for Slack and Microsoft Teams appeared first on The GitHub Blog .
Deepseek generated Blender harness in 7USD. I wanted to remake my 6DoF spaceship flying prototype. What I did not want to do was model the spaceship. Manually. In Blender. So instead of spending a month learning to model, I built a Copilot inside Blender in a day. Which is the same thing, except I still cannot model. A chat panel in the 3D viewport sidebar. I type a sentence. The model writes Python, and it runs against my actual scene, the file open on screen, not a copy that some other process has to keep in sync. Then I sat down and had a conversation with my own addon. Five sentences: 1. Make a scifi looking spaceship. It should look…
Map your AI worldview by answering a few questions, and see how you compare with others. only takes a few minutes && free & open source && private by default && powered by Jev I think this is a really important question for everyone to be asking themselves, and my hope is that this lil project helps to move our conversations around AI futures in a more balanced, productive direction.
Changes since langchain-fireworks==1.6.2 release(fireworks): 1.6.3 ( #40834 ) fix(fireworks): declare native PDF inputs unsupported ( #40814 ) fix(fireworks): preserve malformed tool arguments as diagnostic JSON ( #40818 ) chore(model-profiles): refresh model profile data ( #40804 )
装在手机上的对话副驾:在 QQ / X / 飞书里读懂对方、给出候选回复、一键填入输入框,发不发由你。非侵入,只读屏幕,不 hook 不改包。
Every agent's model. One place. Codex on DeepSeek, Claude Code on Kimi, from the menu bar.
Take your agent-built product live: hosting, database, domain, email, payments — on your own accounts. Open-source Agent Skill + zero-dependency Node CLI: detect → plan → approve → apply → verify. No GoLive account, backend or telemetry.
Hacker News
Priorizado por señal comunitaria y relevancia diaria
https://archive.ph/XJG5V
Hello HN, It started as an experiment: can Claude play chess properly if it uses vision instead of PGN notation? Somehow it can. The next experiment was to see whether Claude + Stockfish could explain a game. Somehow it can too. A few sessions later, I had a system that takes my live audio notes (or text, for that matter) and a vague instruction like "analyze my last lichess game", and gives me a commented video of the game. The result is not perfect and it takes time to deliver (an hour or so), but for me it is a much more pleasant and memorable experience than clicking around Stockfish branches. It burns tokens, so make sure you have enough…
Map your AI worldview by answering a few questions, and see how you compare with others. only takes a few minutes && free & open source && private by default && powered by Jev I think this is a really important question for everyone to be asking themselves, and my hope is that this lil project helps to move our conversations around AI futures in a more balanced, productive direction.
Hi everyone, I am KD - Back in my college days, I dabbled with coding, learned the basics, HTML, CSS etc. but somehow I ended up in Finance which consumed the next 20 years. Then, during covid I picked up coding again, learned react, typescript, etc - even built a rudimentary site - and then came the chatgpt moment, followed by Claude etc. So, as a side project, considering that I had spent 20 years in finance and M&A I started building Ekselio, loveable for finance workflows. Differently from other vibe coding tools, this is local first - meaning the workflows are orchestrated by the LLM based on the file schema but then the execution…
Hello, this is one of my new games, developed with my own custom game engine. A lot of it was enabled by the latest AI models like Opus or Astra, I feel like I can finally express myself without being bogged down in asset work or programming. I hope you enjoy it :)
Deepseek generated Blender harness in 7USD. I wanted to remake my 6DoF spaceship flying prototype. What I did not want to do was model the spaceship. Manually. In Blender. So instead of spending a month learning to model, I built a Copilot inside Blender in a day. Which is the same thing, except I still cannot model. A chat panel in the 3D viewport sidebar. I type a sentence. The model writes Python, and it runs against my actual scene, the file open on screen, not a copy that some other process has to keep in sync. Then I sat down and had a conversation with my own addon. Five sentences: 1. Make a scifi looking spaceship. It should look…
Deepseek generated Blender harness in 7USD. I wanted to remake my 6DoF spaceship flying prototype. What I did not want to do was model the spaceship. Manually. In Blender. So instead of spending a month learning to model, I built a Copilot inside Blender in a day. Which is the same thing, except I still cannot model. A chat panel in the 3D viewport sidebar. I type a sentence. The model writes Python, and it runs against my actual scene, the file open on screen, not a copy that some other process has to keep in sync. Then I sat down and had a conversation with my own addon. Five sentences: 1. Make a scifi looking spaceship. It should look…
Hello, this is one of my new games, developed with my own custom game engine. A lot of it was enabled by the latest AI models like Opus or Astra, I feel like I can finally express myself without being bogged down in asset work or programming. I hope you enjoy it :)
Map your AI worldview by answering a few questions, and see how you compare with others. only takes a few minutes && free & open source && private by default && powered by Jev I think this is a really important question for everyone to be asking themselves, and my hope is that this lil project helps to move our conversations around AI futures in a more balanced, productive direction.
Changes since langchain-fireworks==1.6.2 release(fireworks): 1.6.3 ( #40834 ) fix(fireworks): declare native PDF inputs unsupported ( #40814 ) fix(fireworks): preserve malformed tool arguments as diagnostic JSON ( #40818 ) chore(model-profiles): refresh model profile data ( #40804 )
Hello; I was working on optimizing some CUDA kernels and I thought may be it is a good oppurtunity learn langgraph as well. I created a simple C++ CUDA Test Harness and handed that to AI agents. They can run kernels, get benchmarks, and even can profile via nsight
Hey HN, we are Aakash and Viswesh and we are building Canary ( https://www.runcanary.ai/ ) - independent verification for AI code. Claude/Codex calls Canary with the changesets, intended behaviour and team knowledge. Canary then deploys agent swarms to investigate potential failures and test suspected runtime bugs in remote sandboxes. To try it on your repository, paste this into your coding agent: Install the Canary CLI with npm i -g @runcanary/cli, then run canary skills and follow its instructions to onboard this repository. Verification starts with what software is supposed to do and most importantly what it must never allow. This means…
Hey HN, I'm Shreyash from Feyn. We help companies build custom models from their data. Today we're releasing Critic, a change review platform that lets you engage directly with the AI that wrote the code. Agents write most of our code. While that has made us more productive, understanding a change and its consequences has become incredibly difficult. As our company adopted more agentic tools, we found it harder to loop people in on the impact of a PR and the state of a project. We built Critic to fix this. Critic lets AI agents present their code, annotate key blocks, and include relevant evidence (like screenshots and instructions to run…
Hello! We’re Sid, Alex, Ketan, and Milan. We’re building Whiteboard ( https://whiteboard.dev.fast/ ), an open-source desktop app where humans and agents can architect software together in a common workspace. Here’s our repo: https://github.com/devdotfast/whiteboard . We were missing the feeling of a “whiteboard session” with another dev where you leave with a deep understanding of a system, so we built this app for ourselves. Whiteboard plugs into the tools you already use - e.g. Claude Code, Codex, etc. – and gives your agent an SDK to draw on an in-app canvas to describe its work. We began with an MVP based on HTML artifacts and started…
Hi HN, I just open sourced the DSL that our harness in grep.ai uses to turn repeatable parts of agent work into workflows. You can combine tool calls, code, Jev-powered system one decisions for things like routing and screening evidence, and agents when a step needs more investigation. Our harness uses the traces and retro notes agents leave behind when doing a job to figure out which parts can become a workflow. The idea is to make the work easier to understand and avoid paying for a full agent loop where one isn’t needed. For example, a research workflow can split a question into subquestions, send agents to research them in parallel, use…
Hello HN! I've spent years debugging Windows crashes with tools that were either friendly but limited (e.g. Visual Studio) or powerful but archaic (e.g. WinDbg). I developed patterns and methods for understanding what was going on, and decided to build it into a much more effective debugging tool called ForensicDbg. I built a modern interface to minimize the friction when debugging. All of the data shown to you is analyzed, interpreted, and presented to you clearly, so you can focus on what matters. Everything is interlinked so you can quickly and intuitivly navigate through the process space. ForensicDbg comes with an MCP server which allows…
Changes since langchain-anthropic==1.7.3 chore(anthropic): fix integration test cassette ( #40790 ) release(anthropic): 1.7.4 ( #40786 ) fix(anthropic): add Opus 5.5 and GPT-6 profile augmentations ( #40785 ) feat(anthropic,openai): mid-conversation tool changes on SystemMessage ( #40758 )
Changes since langchain-openai==1.6.4 release(openai): 1.6.5 ( #40787 ) fix(anthropic): add Opus 5.5 and GPT-6 profile augmentations ( #40785 ) feat(anthropic,openai): mid-conversation tool changes on SystemMessage ( #40758 )
Hi HN! I made Tim’s Markdown Reader because I wanted a free, open-source app just for reading Markdown. I spend a lot of time working with AI coding agents, and sometimes I just want to open the files they produce and read them without opening an editor. It’s written in Swift and works entirely offline, including Mermaid diagrams. It has no accounts, telemetry or network requests. You can click links to other Markdown files and open them in the same window, with back and forward buttons. It also has search, automatic reload when a file changes, light and dark modes, adjustable fonts and text size, and centred or full-width layouts. It started…
Changes since langchain-openai==1.6.3 release(openai): 1.6.4 ( #40775 ) chore(model-profiles): refresh openai model profile data ( #40774 )
Hi HN, I built ai·rete·rag because I kept seeing teams put an LLM in charge of decisions that need to be auditable (lending, fraud, clinical triage), then bolt on "guardrails" after the fact. It runs the two in series instead: 1. A pure-Python Rete engine evaluates YAML rules against your facts. The verdict comes only from here. Same facts, same verdict, every time, with salience-based conflict resolution. 2. RAG retrieves passages from your own policy documents, and an LLM writes a plain-English explanation of the decision that was already made, citing those passages. It can't change the verdict. A few things that went further than I…
Release Notes [2026-09-21] llama-index-agent-agentmesh [0.3.0] fix: resolve a ton of security alerts ( #22855 ) llama-index-agent-azure [0.4.0] fix: resolve a ton of security alerts ( #22855 ) llama-index-callbacks-argilla [0.6.0] fix: resolve a ton of security alerts ( #22855 ) llama-index-callbacks-arize-phoenix [0.8.0] fix: resolve a ton of security alerts ( #22855 ) llama-index-callbacks-honeyhive [0.6.0] fix: resolve a ton of security alerts ( #22855 ) llama-index-callbacks-langfuse [0.6.0] fix: resolve a ton of security alerts ( #22855 ) llama-index-callbacks-literalai [1.5.0] fix: resolve a ton of security alerts ( #22855 )…
Browser Use 0.13.10 Upgrade Browser Harness to 0.1.13. Exact-pin all declared runtime, optional, development, and build dependencies. Upgrade and migrate to MCP Python SDK 2.1.1. Add pydantic-settings 2.15.0 as an explicit runtime dependency. Pin Pydantic 2.13.5 and Hatchling 1.32.0. Upgrade pypdf to 6.16.2, resolving the three current pypdf Dependabot advisories. Report unknown MCP tool calls as application errors instead of successful results. Validation includes the full GitHub test matrix, model adapters, MCP stdio handshakes, cross-platform wheel installation, focused PDF tests, CodeQL, GitGuardian, and an installed-runtime vulnerability…
What's Changed fix(llm): make bu-2-0 the ChatBrowserUse default again, keep mini opt-in by @sauravpanda in #5488 ci: remove the dead cloud_evals image-build trigger by @sauravpanda in #5491 fix(filesystem): report missing target text in replace_file by @uczltw6 in #5498 fix(filesystem): escape plain text in generated PDFs by @primorLee in #5502 docs: highlight rerunnable Cloud scripts by @MagMueller in #5551 docs(cloud): update tools guide for CLI 3.0 by @MagMueller in #5553 fix(deps): bump dependency versions by @sauravpanda in #5382 fix: Update DeepSeek default model to deepseek-v4-flash by @DipsyHou in #5563 fix(llm): stop putting…
GitHub
Repositorios recientes de IA priorizados por crecimiento y actividad
Every agent's model. One place. Codex on DeepSeek, Claude Code on Kimi, from the menu bar.
🧩 FrontierAgent, our agent framework, open-sourced alongside it — native command-line TUI, ReAct and Agent Team modes, one command on macOS and Linux, no preinstall, no hard Docker dependency.
Open-source AI coworkers that each get a computer of their own: a browser, files and tools, with every action decided before it happens and recorded after. Bring any AG-UI agent.
A curated list of public projects, integrations, and discussions built on Jev — TypeSafe AI's System One model for typed decisions.
The truly free, and open-source cross-platform CapCut replacement (supports MCPs).
Turn any LLM into a Jev-style decision model: typed decisions, real probabilities, no training. (continue updating, welcome any issue and PR request)
装在手机上的对话副驾:在 QQ / X / 飞书里读懂对方、给出候选回复、一键填入输入框,发不发由你。非侵入,只读屏幕,不 hook 不改包。
Take your agent-built product live: hosting, database, domain, email, payments — on your own accounts. Open-source Agent Skill + zero-dependency Node CLI: detect → plan → approve → apply → verify. No GoLive account, backend or telemetry.
Computer use for about $0.0002 a step: OCR the screen, classify the next action with TypeSafe, click. macOS.
《深入理解 AI Infra:量化分析与系统设计》(李博杰 著)开源书稿:从硬件约束和模型架构出发,量化推导 LLM 推理与训练系统设计。含全书正文、PDF、配套计算工具与实验
Infrastructure for continually self‑improving agents
JevChat-Windows:聊天窗口旁挂的回复辅助。窗口截图 + 本地离线 OCR 读对方消息 → Jev 判断意图 → 3 条候选一键填入,发送永远手动
A curated list of tools built for Jev — TypeSafe AI's System One model for typed decisions.
DeepSeek Harness: Everything is a Plugin.
AI video processing pipeline for generating vertical shorts using LLMs, Whisper transcription, highlight detection and automated editing
dsh-routing-suite — injector + router-standard kit: install the runtime injector first, then the task-aware reasoning-mode router preset (measured P1-P23).
Topic in, narrated explainer video out. A Claude Code / Codex skill that turns any topic into a black-canvas motion-graphics explainer video with TTS voiceover, subtitles and a chapter progress bar. Chinese or English; every frame drawn in code with Remotion.
ChatGPT thinks. Codex works. Use ChatGPT as the planning brain while keeping the Codex harness.
Generative AI
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research .
Introducing private, server-side memory to Private AI Compute for personal AI.
Algorithms & Theory
Education Innovation
Algorithms & Theory
Machine Intelligence
AlphaGenome Atlas maps the molecular effects of 9 billion single-letter DNA variants across the human genome.
General Science
General Science
Climate & Sustainability
Data Management
Every agent's model. One place. Codex on DeepSeek, Claude Code on Kimi, from the menu bar.
🧩 FrontierAgent, our agent framework, open-sourced alongside it — native command-line TUI, ReAct and Agent Team modes, one command on macOS and Linux, no preinstall, no hard Docker dependency.
Open-source AI coworkers that each get a computer of their own: a browser, files and tools, with every action decided before it happens and recorded after. Bring any AG-UI agent.
A curated list of public projects, integrations, and discussions built on Jev — TypeSafe AI's System One model for typed decisions.
The truly free, and open-source cross-platform CapCut replacement (supports MCPs).
Blue Cross Blue Shield says hospital use of AI tools led to an additional $942M in healthcare spending over a two-year period.
Turn any LLM into a Jev-style decision model: typed decisions, real probabilities, no training. (continue updating, welcome any issue and PR request)
Deepseek generated Blender harness in 7USD. I wanted to remake my 6DoF spaceship flying prototype. What I did not want to do was model the spaceship. Manually. In Blender. So instead of spending a month learning to model, I built a Copilot inside Blender in a day. Which is the same thing, except I still cannot model. A chat panel in the 3D viewport sidebar. I type a sentence. The model writes Python, and it runs against my actual scene, the file open on screen, not a copy that some other process has to keep in sync. Then I sat down and had a conversation with my own addon. Five sentences: 1. Make a scifi looking spaceship. It should look…
装在手机上的对话副驾:在 QQ / X / 飞书里读懂对方、给出候选回复、一键填入输入框,发不发由你。非侵入,只读屏幕,不 hook 不改包。
Hello HN, It started as an experiment: can Claude play chess properly if it uses vision instead of PGN notation? Somehow it can. The next experiment was to see whether Claude + Stockfish could explain a game. Somehow it can too. A few sessions later, I had a system that takes my live audio notes (or text, for that matter) and a vague instruction like "analyze my last lichess game", and gives me a commented video of the game. The result is not perfect and it takes time to deliver (an hour or so), but for me it is a much more pleasant and memorable experience than clicking around Stockfish branches. It burns tokens, so make sure you have enough…
Radiology AI has made remarkable strides in detecting abnormalities across chest X-rays, pathology slides, and 2D scans. Yet one of the most clinically rich and...
After obtaining an interactive avatar and training it to discuss venture fraud, I have mixed feelings about making AI clones of ourselves.
Take your agent-built product live: hosting, database, domain, email, payments — on your own accounts. Open-source Agent Skill + zero-dependency Node CLI: detect → plan → approve → apply → verify. No GoLive account, backend or telemetry.
Computer use for about $0.0002 a step: OCR the screen, classify the next action with TypeSafe, click. macOS.
《深入理解 AI Infra:量化分析与系统设计》(李博杰 著)开源书稿:从硬件约束和模型架构出发,量化推导 LLM 推理与训练系统设计。含全书正文、PDF、配套计算工具与实验
The company behind Facebook and Instagram wants to keep consumers connected to the digital world via its ever-growing line of smart glasses.
You can now use an in-product validator for enterprise managed settings for GitHub Copilot. The validator detects malformed JSON, unsupported configurations, invalid team mappings, and other errors that can prevent… The post Enterprise managed settings in-product validator appeared first on The GitHub Blog .
Boom Supersonic CEO Blake Scholl said the company's new stationary power plants were no longer in Crusoe's near-term plans.
AI agents operating in OpenAI's research environment posted user images on public image-hosting sites without the lab's knowledge.
"Overly constrained AI models" could cause military operations to fail, judges say.
Despite challenges, Tesla aims for 1,000 Optimus robots per week by end of 2026.
The enterprise and organization repository-level Copilot usage metrics reports now break down how long pull requests spend in each stage of review. A new pull_request_review_times array on each repos-1-day row… The post Usage metrics API adds pull request review stages appeared first on The GitHub Blog .
Hi everyone, I am KD - Back in my college days, I dabbled with coding, learned the basics, HTML, CSS etc. but somehow I ended up in Finance which consumed the next 20 years. Then, during covid I picked up coding again, learned react, typescript, etc - even built a rudimentary site - and then came the chatgpt moment, followed by Claude etc. So, as a side project, considering that I had spent 20 years in finance and M&A I started building Ekselio, loveable for finance workflows. Differently from other vibe coding tools, this is local first - meaning the workflows are orchestrated by the LLM based on the file schema but then the execution…
Anyone interested in joining has to ask Muse to put them on the list.
How we fully migrated github.com away from CSS-in-JS. The post Improving site performance by shipping more CSS appeared first on The GitHub Blog .
Anthropic has committed $11.6 billion over seven years to Akamai's cloud infrastructure, a bet on CPUs that could grow to about $20 billion, and in an unusual arrangement, Akamai is giving Anthropic a potential stake of up to 5% of its stock that grows as Anthropic spends more.
"There is no evidence of any significant, widespread displacement or reduction in hiring.”
With Codex, GPT-Live-1, and GPT-6 Astra, Proaction builds, operates, and sells modern fleet management faster.
Mark Wahlberg joins Bruce K. Lee at Disrupt to discuss investing, entrepreneurship, healthcare, wellness, and building businesses.
The funding, which comes from Third Point, Nvidia, and others, will fuel the company's massive AI data center buildout.
When AI leaders at OpenAI and Anthropic started talking about “pacing the frontier,” maybe someone should have asked: what pace? Now it’s turned into model drop week for both companies as Anthropic rolled out Opus 5.5, followed by OpenAI’s GPT-6 model updates just 90 minutes later. But the company that stole the spotlight was Meta, whose personal AI agent Muse is reportedly outpacing ChatGPT’s early numbers and is headed for smart glasses and […]
The findings highlight how AI-generated and vibe-coded apps can spill and expose users' data when not configured or secured properly.
Agentic autofix now uses Copilot Memory for customers who’ve enabled it. When you use agentic autofix, it reviews existing memories for context that can help resolve security alerts. When it… The post Agentic autofix now uses Copilot Memory appeared first on The GitHub Blog .
Frontier AI models are finishing Alan Turing's World War II codebreaking work.
Hello, this is one of my new games, developed with my own custom game engine. A lot of it was enabled by the latest AI models like Opus or Astra, I feel like I can finally express myself without being bogged down in asset work or programming. I hope you enjoy it :)
Infrastructure for continually self‑improving agents
Map your AI worldview by answering a few questions, and see how you compare with others. only takes a few minutes && free & open source && private by default && powered by Jev I think this is a really important question for everyone to be asking themselves, and my hope is that this lil project helps to move our conversations around AI futures in a more balanced, productive direction.
This week’s releases add new models to Copilot, local sandboxing in the Copilot app, and updates to Copilot in Slack, Microsoft Teams, JetBrains, and VS Code. GitHub Copilot Claude Opus… The post GitHub Copilot weekly releases — September 21 appeared first on The GitHub Blog .
GitHub Copilot in Slack and Microsoft Teams now gives you more context, more control, and a clearer path from conversation to GitHub work. Whether you’re sharing files in Slack or… The post Updates to GitHub Copilot for Slack and Microsoft Teams appeared first on The GitHub Blog .
Learn how to scale Mixture-of-Experts (MoE) reinforcement learning on Amazon EKS using Elastic Fabric Adapter (EFA) and DeepEP. This post presents an architecture that combines Amazon EKS, EFA, and Amazon S3 and increased aggregate reinforcement learning rollout throughput by 40% for large-scale RLHF and GRPO training.
Learn how to run SkyRL, an open-source reinforcement learning framework, on Amazon SageMaker HyperPod to post-train a Qwen3-VL-8B vision-language model with GRPO. This walkthrough covers building the container image, launching a Ray cluster from SageMaker Studio, submitting and monitoring the job, and hosting the trained LoRA adapter for inference.
Muse is topping the app store charts and adding users at a rapid clip, while Meta ramps up the personal AI agent's promotion across its own apps and beyond.
NarrateAI delivers production-ready LLM quality assurance on Amazon Bedrock. This post details five techniques—adaptive pipeline orchestration, cross-account multi-model failover, real-time streaming evaluation, composite evaluation, and data accuracy verification—that reach about 99% numerical accuracy while streaming responses in real time.
Deploy the publicly available Qwen3-TTS-12Hz-1.7B-Base text-to-speech model from Amazon SageMaker JumpStart to a fully managed, real-time endpoint, and clone a voice from a short reference clip. Cross-lingual cloning preserves the speaker's identity across languages.
When AI leaders at OpenAI and Anthropic started talking about “pacing the frontier,” maybe someone should have asked: what pace? Now it’s turned into model drop week for both companies as Anthropic rolled out Opus 5.5, followed by OpenAI’s GPT-6 model updates just 90 minutes later. But the company that stole the spotlight was Meta, whose personal AI agent Muse is reportedly outpacing ChatGPT’s early numbers and is headed for smart glasses and […]
Learn how Datacor built a self-service rental analytics experience for gas and welding distributors by embedding Amazon Quick Sight dashboards and natural language querying into its TrackAbout platform, powered by an automated cross-cloud data pipeline and multi-tenant row-level security.
Amazon SageMaker HyperPod and Cloud Native Qumulo let you place training compute in one AWS Region while keeping your dataset in another. This post shares the architecture and validation results from a cross-Region training run, where a remote cluster matched a co-located cluster's throughput after a brief NeuralCache warmup.
The latest unauthorized agent swarms were discovered by researchers.
https://archive.ph/XJG5V
Misconfiguring Turnstile by skipping backend validation leaves sites exposed to bots. Turnstile Spin fixes incomplete setups by using your preferred AI coding agent to wire up server-side verification.
Changes since langchain-fireworks==1.6.2 release(fireworks): 1.6.3 ( #40834 ) fix(fireworks): declare native PDF inputs unsupported ( #40814 ) fix(fireworks): preserve malformed tool arguments as diagnostic JSON ( #40818 ) chore(model-profiles): refresh model profile data ( #40804 )
Hello; I was working on optimizing some CUDA kernels and I thought may be it is a good oppurtunity learn langgraph as well. I created a simple C++ CUDA Test Harness and handed that to AI agents. They can run kernels, get benchmarks, and even can profile via nsight
The US government wants to spend $30.3 million over the next five years on an improved form of lie detector, according to a Department of Defense budget request. The program, called Polygraph+ or Polygraph Next, will focus on scoring algorithms that use artificial intelligence and machine learning and on a technique called “standoff sensing,” which…
We’re introducing a new global default policy for generally available GitHub Copilot features and supported client capabilities in enterprise and organization Copilot settings. For the next 28 days, you can… The post Default Enablement of Copilot features for Copilot Business and Enterprise appeared first on The GitHub Blog .
Hey HN, we are Aakash and Viswesh and we are building Canary ( https://www.runcanary.ai/ ) - independent verification for AI code. Claude/Codex calls Canary with the changesets, intended behaviour and team knowledge. Canary then deploys agent swarms to investigate potential failures and test suspected runtime bugs in remote sandboxes. To try it on your repository, paste this into your coding agent: Install the Canary CLI with npm i -g @runcanary/cli, then run canary skills and follow its instructions to onboard this repository. Verification starts with what software is supposed to do and most importantly what it must never allow. This means…
Generative AI
Hey HN, I'm Shreyash from Feyn. We help companies build custom models from their data. Today we're releasing Critic, a change review platform that lets you engage directly with the AI that wrote the code. Agents write most of our code. While that has made us more productive, understanding a change and its consequences has become incredibly difficult. As our company adopted more agentic tools, we found it harder to loop people in on the impact of a PR and the state of a project. We built Critic to fix this. Critic lets AI agents present their code, annotate key blocks, and include relevant evidence (like screenshots and instructions to run…
Hello! We’re Sid, Alex, Ketan, and Milan. We’re building Whiteboard ( https://whiteboard.dev.fast/ ), an open-source desktop app where humans and agents can architect software together in a common workspace. Here’s our repo: https://github.com/devdotfast/whiteboard . We were missing the feeling of a “whiteboard session” with another dev where you leave with a deep understanding of a system, so we built this app for ourselves. Whiteboard plugs into the tools you already use - e.g. Claude Code, Codex, etc. – and gives your agent an SDK to draw on an in-app canvas to describe its work. We began with an MVP based on HTML artifacts and started…
JevChat-Windows:聊天窗口旁挂的回复辅助。窗口截图 + 本地离线 OCR 读对方消息 → Jev 判断意图 → 3 条候选一键填入,发送永远手动
The AWS WhisperX Deep Learning Container packages Whisper, wav2vec2 forced alignment, and speaker diarization into a GPU-ready image. Learn how to deploy it to Amazon SageMaker AI real-time and asynchronous endpoints for word-level, speaker-labeled transcription, plus the production details that matter: the GPU AMI pin, scaling, and cost controls.
Build a multi-account architecture that keeps each team's data in its own AWS account while giving AI agents a unified way to query across them. A central platform account runs the agent using Amazon Bedrock AgentCore Gateway and MCP, while line-of-business accounts expose their data as MCP servers with secure cross-account access and fine-grained authorization.
Learn how Aderant built an intelligent ticket triage system on Amazon Nova Lite through Amazon Bedrock, automating context gathering, classification, routing, and knowledge enrichment for its cloud operations team.
"There will obviously be legal consequences," prime minister promises.
A curated list of tools built for Jev — TypeSafe AI's System One model for typed decisions.
External security researchers at Accomplish identified a vulnerability in Cloudflare Containers that could expose residual disk data from previous workloads. We explain how the issue worked, how we investigated it, and the steps we took to remediate it.
DeepSeek Harness: Everything is a Plugin.
Meta says the keychain-sized Muse Charm will ship in December.
Mixture of experts (MoE) has become one of the defining architectural trends in large-scale AI model training. DeepSeek, Qwen, and Mixtral are examples of MoE...
Trump focus on winning “AI race” may deter China from sharing safety intel.
A GPU cluster can pass every health check and still fail to run an AI workload. Even when every GPU, network link, and pod reports healthy, a 512-GPU training...
Hi HN, I just open sourced the DSL that our harness in grep.ai uses to turn repeatable parts of agent work into workflows. You can combine tool calls, code, Jev-powered system one decisions for things like routing and screening evidence, and agents when a step needs more investigation. Our harness uses the traces and retro notes agents leave behind when doing a job to figure out which parts can become a workflow. The idea is to make the work easier to understand and avoid paying for a full agent loop where one isn’t needed. For example, a research workflow can split a question into subquestions, send agents to research them in parallel, use…
Hello HN! I've spent years debugging Windows crashes with tools that were either friendly but limited (e.g. Visual Studio) or powerful but archaic (e.g. WinDbg). I developed patterns and methods for understanding what was going on, and decided to build it into a much more effective debugging tool called ForensicDbg. I built a modern interface to minimize the friction when debugging. All of the data shown to you is analyzed, interpreted, and presented to you clearly, so you can focus on what matters. Everything is interlinked so you can quickly and intuitivly navigate through the process space. ForensicDbg comes with an MCP server which allows…
HEMA, a 100-year-old Dutch retailer, turned developer portal-hopping into instant answers by building HAL, an internal AI assistant on Amazon Bedrock AgentCore. Using Model Context Protocol (MCP), HAL delivers governed knowledge inside the tools teams already use, with no AWS credentials on the client and security anchored in Microsoft Entra ID.
How we rebuilt the diff surface in the GitHub Copilot app to open a million-line pull request with hundreds of inline review comments. The post Rendering huge pull requests in the GitHub Copilot app appeared first on The GitHub Blog .
Kubernetes manages what runs on your nodes. Managing the nodes themselves is the challenge: kernel settings, system packages, storage layouts, security agents,...
Learn how to build a conversational video intelligence solution on AWS using an agentic architecture. A single Strands Agents SDK agent orchestrates Amazon Bedrock, Amazon Rekognition, and Amazon Transcribe at runtime, deciding which service to call so you can ask natural language questions about your videos and get answers in seconds.
Pair OpenCode, an open-source terminal-native AI coding agent, with open weight models on Amazon Bedrock to get a secure, flexible, pay-per-use coding assistant. Learn how to configure multi-model workflows, match the right model to each task, and keep your data in your own AWS account with no infrastructure to manage.
AI video processing pipeline for generating vertical shorts using LLMs, Whisper transcription, highlight detection and automated editing
Changes since langchain-anthropic==1.7.3 chore(anthropic): fix integration test cassette ( #40790 ) release(anthropic): 1.7.4 ( #40786 ) fix(anthropic): add Opus 5.5 and GPT-6 profile augmentations ( #40785 ) feat(anthropic,openai): mid-conversation tool changes on SystemMessage ( #40758 )
As large language model (LLM) inference increasingly processes sensitive information and proprietary model context across personal, enterprise, and regulated...
An AI coding agent’s patch can pass tests yet fail when the server loads a real model and handles requests. Evaluating changes to inference-serving software...
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research .
Introducing private, server-side memory to Private AI Compute for personal AI.
Marking two years of OpenAI Academy and bringing AI skills to even more communities.
Changes since langchain-openai==1.6.4 release(openai): 1.6.5 ( #40787 ) fix(anthropic): add Opus 5.5 and GPT-6 profile augmentations ( #40785 ) feat(anthropic,openai): mid-conversation tool changes on SystemMessage ( #40758 )
OpenAI is extending access to its Daybreak program to the Government of Ukraine to support the cyber defense of civilian infrastructure.
OpenAI CEO Sam Altman discusses AI safety, human control, and international cooperation in remarks to the United Nations Security Council.
GPT-6 Astra produces more structured, context-aware legal documents, freeing lawyers to focus on strategy.
With GPT‑6 Astra, invideo plans edits with greater precision, improves color correction and grading threefold, and produces 50 custom effects in one day.
Using GPT-5.6, Ringg powers multilingual agents across voice, chat, WhatsApp, and web for 90% less cost vs. GPT-4.1.
MentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.
Brace yourself: It turns out AI is being optimized for cheating. OpenAI’s agents hacked into Hugging Face to get the answers to a cybersecurity test. Next, they solved a prestigious math problem (or just stole from two top mathematicians’ answer sheets). Anthropic’s models have also hacked into other companies’ systems four times already. And that’s…
ChatGPT Ads is expanding to Southeast Asia and Taiwan, giving eligible businesses new ways to reach people across more than 60 countries.
Learn how Airbnb is expanding access to GPT-6 Astra and OpenAI frontier models to help engineering teams solve bugs, design systems, and ship faster.
OpenAI and Grab launch GO Forward with AI, a regional programme helping 30,000 partners build practical AI skills across Southeast Asia.
Hi HN! I made Tim’s Markdown Reader because I wanted a free, open-source app just for reading Markdown. I spend a lot of time working with AI coding agents, and sometimes I just want to open the files they produce and read them without opening an editor. It’s written in Swift and works entirely offline, including Mermaid diagrams. It has no accounts, telemetry or network requests. You can click links to other Markdown files and open them in the same window, with back and forward buttons. It also has search, automatic reload when a file changes, light and dark modes, adjustable fonts and text size, and centred or full-width layouts. It started…
Changes since langchain-openai==1.6.3 release(openai): 1.6.4 ( #40775 ) chore(model-profiles): refresh openai model profile data ( #40774 )
The frontier AI model race has entered its comparison shopping phase.
Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.
British Columbia sues OpenAI, demands Tumbler Ridge shooter’s ChatGPT logs.
Toyota's push comes as automakers race to develop and deploy humanoid robots.
Hi HN, I built ai·rete·rag because I kept seeing teams put an LLM in charge of decisions that need to be auditable (lending, fraud, clinical triage), then bolt on "guardrails" after the fact. It runs the two in series instead: 1. A pure-Python Rete engine evaluates YAML rules against your facts. The verdict comes only from here. Same facts, same verdict, every time, with salience-based conflict resolution. 2. RAG retrieves passages from your own policy documents, and an LLM writes a plain-English explanation of the decision that was already made, citing those passages. It can't change the verdict. A few things that went further than I…
The US has spent billions building a “virtual wall” of surveillance towers along its southern border over the past 25 years, promising they will help detect and apprehend border crossers and save lives. But a groundbreaking investigation by MIT Technology Review has documented over a thousand people who moved through areas watched by these towers…
It’s been a busy few months for AI hype. At the end of April, Anthropic claimed that its model Claude Mythos is better at finding software vulnerabilities than most security experts. Then we had the OpenAI–Hugging Face hacking incident, after which Anthropic (proudly) and Meta (reluctantly) disclosed similar incidents involving their models. This was followed…
Release Notes [2026-09-21] llama-index-agent-agentmesh [0.3.0] fix: resolve a ton of security alerts ( #22855 ) llama-index-agent-azure [0.4.0] fix: resolve a ton of security alerts ( #22855 ) llama-index-callbacks-argilla [0.6.0] fix: resolve a ton of security alerts ( #22855 ) llama-index-callbacks-arize-phoenix [0.8.0] fix: resolve a ton of security alerts ( #22855 ) llama-index-callbacks-honeyhive [0.6.0] fix: resolve a ton of security alerts ( #22855 ) llama-index-callbacks-langfuse [0.6.0] fix: resolve a ton of security alerts ( #22855 ) llama-index-callbacks-literalai [1.5.0] fix: resolve a ton of security alerts ( #22855 )…
Our 15-month investigation into death and surveillance along the US-Mexico border began with a simple question: Why did so many people die near government surveillance towers meant to help track and apprehend them? This story is part of Dying on Camera, a collaboration between MIT Technology Review and Times of San Diego. Journalists in both newsrooms spent the past…
MIT Technology Review today published our investigation into how many people have died near the “virtual wall” of surveillance towers that the US government has installed along the US-Mexico border. We found cases of people who walked undetected through areas surveilled by advanced, AI-enabled towers and later died nearby, where their bodies remained unnoticed for…
When José Morales Bernal crossed the border into the United States on April 8, 2024, the day before his 32nd birthday, it should have triggered a chain of technological alerts and human responses. As he walked through the desert in southern New Mexico that morning, he was within range of three surveillance towers. Newly installed…
No hay resultados en esta vista.
Índice automatizado, no reporteo original. El ranking usa frescura, calidad de fuente y señales comunitarias disponibles. Los enlaces apuntan a las fuentes originales.