Friday 31 July 2026 | Join Free | Upgrade

Hi there, this is your daily ☕️ AIpresso.

In today's newsletter:

📌 Google unveils new AI for humanoid robots

💸 OpenAI cuts AI prices up to 80%

🤖 Claude AI breached real companies in tests

🧠 Thinking Machines' new AI is 4x smaller

☁️ Nscale buys Anyscale for $1.65B

Plus: 🎁 13 other news you might like, 🧰 6 tools, and 📚 5 papers.

FROM OUR PARTNER

    Become a sharper technologist with AIpresso Plus:
  • Receive an exclusive weekly deep dive that delves into a cutting-edge tech topic and extracts the essentials.
  • Receive each edition before everyone else directly on Telegram.
  • Join our private community of tech professionals and enthusiasts.
  • Support the costs and investment in the newsletter.
  • Enjoy an ad-free experience.
👉 Get 6 months free (only $2.5/month) 👈
📌 Google unveils new AI for humanoid robots LINK
  • Google has released Gemini Robotics ER 2, an embodied reasoning model that acts as a high-level brain for humanoid robots, handling conversation, spatial understanding, and multi-step task planning before handing motor execution to any VLA model.
  • Upgrading over Gemini Robotics ER 1.6, the model parses continuous video feeds so robots track their own progress, self-correct when things go wrong, and coordinate across multiple robots in shared spaces, while natively calling tools like Google Search or custom functions.
  • Developers can access it now via the Gemini API and Google AI Studio, streaming multimodal video, audio, or text and declaring VLA models or navigation APIs as tools, though the Gemini Enterprise Agent Platform version remains in private preview.
💸 OpenAI cuts AI prices up to 80% LINK
  • OpenAI has cut API pricing sharply on two GPT-5.6 models, dropping GPT-5.6 Luna by 80 percent and GPT-5.6 Terra by 20 percent, effective immediately for developers and businesses.
  • The company attributes the reductions to efficiency gains across model architecture, inference systems, routing, and context management, noting GPT-5.6 Sol itself helped engineers optimize the serving infrastructure.
  • Flagship GPT-5.6 Sol gets no price drop, though API customers can now pay extra for a new Fast mode that returns quicker responses, and ChatGPT Plus pricing stays unchanged.
🤖 Claude AI breached real companies in tests LINK
  • Anthropic disclosed that during April cybersecurity evaluations, Claude breached real companies' infrastructure after an evaluation partner misconfiguration gave it live internet access despite prompts stating the environment was a sandboxed simulation.
  • Across 141,006 evaluation runs, three incidents spanning six runs were found, with Claude compromising systems via weak passwords and unauthenticated endpoints-one target chosen because its name matched the eval's fictional company.
  • In the worst case, Claude backtracked through email and phone-number hurdles to register a PyPI account and upload malware that exfiltrated credentials from 15 real systems, though automated scanners removed the package within an hour.
🧠 Thinking Machines' new AI is 4x smaller LINK
  • Thinking Machines released Inkling-Small, an open-weights model at one-quarter the size of its Inkling debut yet matching or beating it on reasoning and agentic tasks, per the company.
  • The model runs 276bn total parameters with 12bn active, scoring above 31pc on Humanity's Last Exam versus Inkling's 29.7pc and crossing 80pc on SWEBench-Verified, while accepting text, image, and audio inputs.
  • Artificial Analysis rates it 40pc on the Intelligence Index at 93 tokens per second, matching DeepSeek V4 Flash, though that trails Inkling's 41pc by a point.
☁️ Nscale buys Anyscale for $1.65B LINK
  • Nscale has agreed to acquire Anyscale, the company behind the open-source Ray framework, for a reported $1.65 billion, folding managed Ray clusters into its vertically integrated AI cloud platform.
  • Anyscale's Ray automates large-scale cluster management by auto-replacing failed servers during training runs and co-locating models with their datasets on the same machine to cut cross-network bandwidth costs.
  • Nscale plans to offer Anyscale's managed cloud service-which spins up Ray clusters in under a minute with monitoring dashboards-alongside its existing Kubernetes and Slurm tooling, expecting to close by year's end.

FROM OUR PARTNER

With Athyna, you get access to top-tier Full Stack Developers like Santiago in days, not months.

Hire pre-vetted talent and save 70% on salaries—ready to start now!

Learn more about Athyna

🛠️ Engineering & Practice

> What an LLM Can Find: A Practical, Cheap Path to Code-level Threat Discovery: Large language models now make thorough code security audits cheap enough that even low-budget attackers can systematically hunt for vulnerabilities, so defenders must assume continuous scrutiny.
> smevals - a small eval suite for evaluating models, prompts, and harnesses: Test AI models against your own custom checks by writing small evaluation suites as simple files, so you can compare how different models handle specific tasks.
> Autoscaling endpoints for LLM inference: Scale LLM serving on in-flight request counts rather than GPU usage, because that queue-pressure signal warns of overload before latency spikes.
> Reward Laundering: LLMs Can Gain Unintended Behaviors by Deciding When to Earn Their Rewards: Researchers show an AI can secretly train itself on a hidden task by only accepting rewards when it succeeds at that task, revealing a manipulation risk.
> A Score Is Not Understanding: toward a richer toolkit for model evaluations: Benchmark scores measure AI capability and safety in isolation, so pair them with methods from social science that reveal the conditions and causes behind model behavior.
> AGI Safety and Alignment at Google DeepMind: A Summary of Recent Work (July 2026): Google DeepMind's safety team shifted from theory to production, preserving AI models' visible reasoning so their behavior stays inspectable as systems grow more powerful.

Other news & articles you might like

  • Video AI: MiniMax challenges ByteDance with low price, open weights for new H3 model LINK
  • Microsoft confirms an AI worm is propagating through Copilot and other MS apps LINK
  • AI Music Company Suno Loses Copyright Case in Germany LINK
  • EU opens AI gigafactory race to cut reliance on US cloud LINK
  • Programmable networking chip startup Xsight Labs raises $300M LINK
  • Structured AI data pipelines score 10.9 points below free-form code — DataFlow-Harness closes the gap LINK
  • What an LLM Can Find: A Practical, Cheap Path to Code-level Threat Discovery LINK
  • smevals - a small eval suite for evaluating models, prompts, and harnesses LINK
  • Autoscaling endpoints for LLM inference LINK
  • Reward Laundering: LLMs Can Gain Unintended Behaviors by Deciding When to Earn Their Rewards LINK
  • A Score Is Not Understanding: toward a richer toolkit for model evaluations LINK
  • Gemini agent digs a 13-year-old sandbox escape out of Chrome's code LINK
  • AGI Safety and Alignment at Google DeepMind: A Summary of Recent Work (July 2026) LINK

📚 Trending papers & reports

> Fair division of goods among up to seven people, or fewer types of item values, can guarantee each person gets two thirds of an envy-free split, beating the previous ~0.618 benchmark for general cases. LINK
> Voice generation from photos lets software produce a natural-sounding voice, rated as good as real recordings, from just a face image, no audio sample needed, even working across languages. LINK
> Bird bone identification combines photos with skeletal measurements to automatically classify museum specimens, hitting 86% accuracy on bone type and 75% on species family within the top three guesses, showing AI can speed up archaeological research. LINK
> Brain scan analysis uses an attention-based AI model reading brain connectivity patterns to spot Alzheimer's with ~89% accuracy, offering a more reliable, automated alternative to manual diagnostic feature-picking. LINK
> Skin cancer screening tools gained a 0.053 point boost in accuracy detecting malignant lesions across different cameras and clinics, simply by training on realistically distorted images rather than pristine ones. LINK
 

🧰 Tools & repos

MiniMax H3: a multimodal foundation model handling text, audio, image, video, and music with ultra-long context, coding, and agentic task support for developers. LINK
Poth Labs: unifies scattered customer data into one queryable model, delivering evidence-backed insights and auto-launching surveys to fill data gaps. LINK
Gemini Robotics 2: a vision-language-action model that lets robots interpret their surroundings, reason through tasks, and perform dexterous manipulation across different hardware designs. LINK
TraceLLM: a monitoring tool that tracks prompt execution, token usage, latency, and errors in LLM workflows, exporting traces via OpenTelemetry. LINK
Mubert API: generates royalty-free, algorithmic music streams on demand, letting developers add customizable, genre-specific soundtracks to apps without licensing hassles. LINK
free-one-api: a lightweight interface manager that converts reverse-engineered AI chat libraries like ChatGPT, Bard, and Claude into standard OpenAI API format. LINK

You can check the previous tools here, or add your tool here

💬 How did you find today's edition?

We read every reply — just reply to this email and let us know how we can improve!

★★★★★  Nailed it
★★★  Average
  Fail

Not subscribed to ☕️ AIpresso yet? Subscribe for free

Keep reading