top of page
Catapult your career and thrive in the ai revolution

AI Insights in Action


01:30
OpenAI Says Its Agents Outwork Its Humans 3 To 1 #aiagents
OpenAI just published its own numbers on how much research its coding agents handle. The headline figure is 3.1. The footnote is where the story lives.
On September 6, 2026, OpenAI posted "Research acceleration: The view inside OpenAI" — an unusually direct look at how its own researchers work day to day. The number everyone grabbed: by mid-August, researchers were running 3.1 agent-workdays for every one human workday across the research organization.
Here is what that figure actually measures. It is aggregate machine runtime — roughly 24.8 agent-hours logged against a single eight-hour human workday, counting agents running concurrently and any subagents they spin up downstream. It is a measure of hours executed, not problems solved. An agent that spends four hours debugging itself, retrying a failed approach, babysitting a training run, or generating code a researcher later throws away still contributes to that total. OpenAI does not claim it produces 3.1 times more research, and the distinction matters more than the number.
The company also says it hit the milestone Sam Altman set out in October 2025: an intern-level AI research assistant by September 2026, on the way to what he described as an automated AI researcher by March 2028. OpenAI defines that intern narrowly — a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days. It does not choose what to work on.
The supervision numbers are the honest part of the post. On successful tasks estimated to take a human four to eight hours, more than half still involved at least one human intervention over the previous six months. Agents work across what OpenAI frames as six stages — Decide, Design, Build, Run, Analyze, Communicate — but the human work has shifted rather than disappeared. Engineers now spend more time reviewing code diffs and deciding what is safe to ship. Supervision, not compute, is emerging as the bottleneck.
Then there is the cost. The median OpenAI researcher was spending over $600 a day on inference. At the 90th percentile, that figure passes $7,000 a day per researcher. Developer Simon Willison, reading the same charts, noted spend per researcher climbed from near zero in February 2026 to roughly $600 by late August, with a sharp jump in late July — his guess being that internal staff got access to the model later released as GPT-6 Astra.
The benchmark picture stays sober too. On PaperBench, which tests whether a system can replicate published research, agents score around 21%. On MLE-bench, roughly 11%. Those are real capabilities and they are nowhere near a researcher.
One more detail: after July outages tied to cybersecurity concerns, OpenAI restricted internal use of its Astra model. Astra's share of GPU allocation fell 59.2% — but total workloads simply moved to other models. Restricting one agent did not reduce the compute; it redistributed it.
Why this matters to you even if you never write a line of code: "3.1 agent-workdays per human workday" is exactly the shape of metric that will land in your workplace next. Seat counts. Hours logged. Tasks assigned to an assistant. All of it measures activity, and none of it measures output. When a vendor or an executive quotes you a multiplier on AI productivity, the question worth asking is whether anyone measured the result or just the runtime — and who is now doing the reviewing that used to be the doing. To its credit, OpenAI published both halves. Most of the coverage only carried one.
Have you seen an AI productivity number at your job that measured activity instead of results? Tell me about it in the comments.
If this was useful, hit like, subscribe, and turn on the bell for more AI news translated into plain English. Share it with whoever sent you the 3.1 headline. And if you want to support the channel directly, the Patreon link is below.
Keywords: OpenAI research acceleration, agent workdays, automated research intern, Sam Altman, AI coding agents, OpenAI agents, AI productivity metrics, recursive self improvement, PaperBench, MLE-bench, GPT-6 Astra, AI research automation, inference cost, AI agent supervision, Jakub Pachocki
👇 Connect with The AI Guide:
Website: www.theaiguide.ai
YouTube: https://www.youtube.com/channel/UCamFJyTjb_kqUbNpVUVwQeg
Patreon: www.patreon.com/theaiguide
Facebook: www.facebook.com/davidtheaiguide
Instagram: www.instagram.com/theaiguide
LinkedIn: https://www.linkedin.com/company/the-ai-guide-on-youtube/
Sources:
1. Simon Willison on OpenAI's post — https://simonwillison.net/2026/Sep/6/research-acceleration-the-view-inside-openai/
2. eWeek — https://www.eweek.com/news/openai-ai-research-intern-agent-workdays/
3. The New Stack — https://thenewstack.io/openai-agent-research-bottleneck/
4. MLQ News analysis — https://mlq.ai/news/openais-31-agent-workday-figure-measures-machine-runtime-not-31-times-more-research/
#AINews #OpenAI #AIAgents #AI #TheAIGuide

09:27
Uber Just Got Fined $1 Billion Over an Algorithm. Then This. #ai
Two Uber drivers. Same car park, same app, same job. One offered £27. The other £23. A claim covering 241,000 drivers says that gap is the point.
On September 2, 2026, Dutch foundation Stichting Worker Info Exchange International filed a collective claim in Amsterdam district court for roughly 241,000 Uber drivers across seven countries: the UK, the Netherlands, Belgium, France, Germany, Poland and Romania. It is led by James Farrar, who won the UK Supreme Court case that reclassified Uber drivers as workers.
Here is the part almost every write-up skipped. This is not a minimum-wage case. It is a data-protection case, run on Article 22 of the GDPR — the right not to be subject to a decision based solely on automated processing that significantly affects you. The allegation: Uber uses automated decision-making, including profiling, to set individual pay and allocate work. What is sought is transparency, not a higher rate.
WHAT THE CLAIMANTS SAY HAPPENED. Kola Oba, 48, from north London, was offered £23 for a job another driver in the same car park was offered £27 for. Oba believes the lower offer followed a run of cheap fares he had accepted, and called the system "soulless": "they have all my information and they are using it against my own wellbeing." That is the claimants' account of one incident as reported by the Guardian — not a finding of fact. The claim also estimates UK drivers have lost roughly £5,000 a year. Claimants' figures, not an audit.
WHAT UBER SAYS. Uber categorically rejects the allegations. The company says fares are calculated from real-time information about the trip itself — journey, duration and destination — that drivers see their earnings and destination before accepting, and that the percentage Uber keeps has stayed flat. Its position is that dynamic pricing raises pay on less desirable trips, rather than adjusting to a driver's history.
THE CONTEXT THAT MAKES THIS MORE THAN A LAWSUIT. Two weeks earlier, on August 23, 2026, the Dutch Data Protection Authority fined Uber nearly €825 million for deactivating drivers' accounts through a fully automated process. Different decision, same country, same law, same article. A regulator has already found that an automated decision about a driver can breach it.
WHY A US VIEWER SHOULD CARE ABOUT A DUTCH FILING. Law professor Veena Dubal set out the case for what she calls algorithmic wage discrimination in the Columbia Law Review, using US platforms; CBS News has covered it here. In an audit of 500 workforce-management AI vendors, Dubal and Wilneida Negrón flagged 20 at high risk of enabling what they call "surveillance pay" — wage-setting built on granular, real-time monitoring data. Those vendors serve healthcare, call centres, logistics and retail. Not just ride-hail.
THE HONEST CLOSE. No court has ruled, Uber denies everything, and the right this case runs on is European and British. There is no federal US equivalent, so an American worker cannot demand the same disclosure. The thing to watch isn't the verdict. It is whether pay stops being a posted number and becomes a personal one, quietly, everywhere.
WHAT TO DO WITH THIS: if any part of your pay moves — commission, bonus, shift rate, per-job rate — ask, in writing, how it is calculated. Not to sue anyone. Not knowing how your own pay is set is the condition this model depends on.
Do you actually know how your own pay is calculated — all of it, not just the base? Tell me in the comments.
If you want AI news with the sources actually checked, like this video, subscribe and hit the bell. Share it with someone whose pay moves week to week, and there's a Patreon link below.
Keywords: algorithmic pay, Uber algorithm lawsuit, GDPR Article 22, automated decision making, algorithmic wage discrimination, surveillance pay, Worker Info Exchange, James Farrar, Veena Dubal, gig economy pay, AI and wages, Uber class action, AI at work, data protection
👇 Connect with The AI Guide:
Website: www.theaiguide.ai
YouTube: https://www.youtube.com/channel/UCamFJyTjb_kqUbNpVUVwQeg
Patreon: www.patreon.com/theaiguide
Facebook: www.facebook.com/davidtheaiguide
Instagram: www.instagram.com/theaiguide
LinkedIn: https://www.linkedin.com/company/the-ai-guide-on-youtube/
🔗 Sources:
[1] The Guardian (2 Sep 2026): https://www.aol.co.uk/articles/uber-drivers-launch-european-class-153726000.html
[2] Personnel Today: https://www.personneltoday.com/hr/uber-collective-action/
[3] Irish Examiner: https://www.irishexaminer.com/world/arid-41905981.html
[4] Dutch DPA: https://www.autoriteitpersoonsgegevens.nl/en/current/uber-fined-nearly-825-million-euros-for-automated-driver-blocking
[5] Equitable Growth: https://equitablegrowth.org/how-artificial-intelligence-uncouples-hard-work-from-fair-wages-through-surveillance-pay-practices-and-how-to-fix-it/
[6] CBS News: https://www.cbsnews.com/news/algorithmic-wage-discrimination-artificial-intelligence/
#ArtificialIntelligence #AIatWork #GigEconomy #AInews #TheAIGuide

01:30
An AI Designed This Drug. Patients Got Younger. #ai
An AI designed the drug. Six separate aging clocks — built by different teams, on different data — all said the patients taking it looked biologically younger.
Here's what actually happened. Insilico Medicine, a biotech that uses generative AI to find drug targets and design molecules, developed a compound called rentosertib. It was never meant to be an anti-aging drug. It was built to treat idiopathic pulmonary fibrosis, a lung disease that scars lung tissue until breathing gets harder. It targets TNIK, a protein Insilico's AI flagged — a gene the company says touches six recognized hallmarks of aging.
The company ran a Phase 2a trial across 21 sites in China. Of the 71 patients enrolled, 42 consented to have their blood profiled over 12 weeks. Researchers tracked 2,841 proteins in those blood samples, then fed the results through six independently built proteomic aging clocks: ProtAge, PAC, ipfP3GPT, PAOPAC, and two variants of OrganAge — one trained to predict chronological age, one trained on mortality risk. They were built with different methods, from classical machine learning to deep learning, and share neither features nor training data.
The result, published in Nature Biotechnology: at week four, patients on the 30 mg twice-daily dose read roughly 2.7 to 3.5 years younger than expected on the chronological clocks. Some mortality-trained models showed larger shifts. Five of the six clocks showed a negative effect size in that arm. Across all six clocks and three post-baseline timepoints, 21 of 54 comparisons reached statistical significance. The effect plateaued by week 12.
On lung function, the 60 mg once-daily group showed a mean improvement of 98.4 mL in forced vital capacity, while placebo declined by 20.3 mL. Notably, the biggest lung improvement and the strongest clock signal came from different dose arms — which the authors argue cuts against the simplest explanation.
And that explanation is the catch. Pulmonary fibrosis is an inflammatory disease, and inflammation shows up in blood proteins. If the drug heals the lung, the blood improves, and the clocks might simply be reading a healthier patient rather than a younger one. The authors say this plainly: full disentanglement of aging effects from disease effects is not achievable inside an IPF cohort. Proving it would take a trial in healthy volunteers. They list the limits themselves — small sample, 12-week window, and an analysis leaning on computation without complementary omics data or direct senescent-cell measurement.
Independent reaction has been careful, not dismissive. Michael Levitt, the 2013 Nobel laureate in chemistry, said what convinced him was not the size of the effect but the agreement between models that share neither features nor training data. He has argued for follow-up work in healthy volunteers.
Why this matters to you even if you never take this drug: this is the first time a drug an AI helped design has been measured against aging biomarkers inside a real clinical trial. No regulator anywhere licenses a medicine based on proteomic age scores — that is not a category that exists yet. But rentosertib entered Phase 3 in July, and the measurement approach used here is the one that would have to be validated before any "biological age" claim could ever reach a label. Expect "biologically younger" to show up in supplement marketing long before it shows up on anything approved. Knowing where the real evidence stops is the useful part.
What would convince you that a drug actually slowed aging rather than just treated a disease? Tell me in the comments.
If this was useful, hit like, subscribe, and turn on the bell so you don't miss the next one. Share it with someone who keeps sending you longevity headlines. And if you want to support the channel directly, the Patreon link is below.
Keywords: rentosertib, Insilico Medicine, AI designed drug, proteomic aging clocks, biological age reversal, TNIK inhibitor, idiopathic pulmonary fibrosis, Nature Biotechnology study, AI drug discovery, longevity research, aging biomarkers, Phase 2a trial, Alex Zhavoronkov, generative AI biotech, anti aging science
👇 Connect with The AI Guide:
Website: www.theaiguide.ai
YouTube: https://www.youtube.com/channel/UCamFJyTjb_kqUbNpVUVwQeg
Patreon: www.patreon.com/theaiguide
Facebook: www.facebook.com/davidtheaiguide
Instagram: www.instagram.com/theaiguide
LinkedIn: https://www.linkedin.com/company/the-ai-guide-on-youtube/
Sources:
1. Insilico Medicine announcement — https://insilico.com/news/rnt0709261-rentosertib-proteomic-aging-clocks
2. Longevity.Technology analysis — https://longevity.technology/news/rentosertib-puts-aging-clocks-to-the-clinical-test/
3. The Next Web — https://thenextweb.com/news/insilico-rentosertib-proteomic-aging-clocks
4. Unite.AI technical breakdown — https://www.unite.ai/proteomic-aging-clocks-track-biological-age-reversal-in-rentosertib-trial/
#AINews #AIDrugDiscovery #Longevity #TheAIGuide

08:07
OpenAI Says GPT-6 Astra Is #1. The Independent Lab Says #2. #ai
OpenAI called GPT-6 Astra the most intelligent model in the world. The independent lab that scores these things put it second. 99% AGI my a** — here's which one matters to you.
GPT-6 Astra reached a limited set of organizations on September 3, 2026, then paid ChatGPT users — Plus, Pro, Business, Enterprise — the next day, plus the API and AWS Bedrock. Roughly a 1,050,000-token context window, knowledge cutoff April 30, 2026. Pricing: $10 per million input tokens, $50 per million output.
THE HEADLINE NUMBERS. OpenAI reported 99.9% on the upgraded ARC-AGI-3 against 7.8% for GPT-5.6 Sol, 97.6% on FrontierMath Tier 4, 96.0% on GPQA Diamond and 100% on ExploitBench. Impressive — and worth reading with the settings in mind: OpenAI ran these at max effort and removed standard time limits on the cyber test.
THE SCORES IN CONTEXT. On the Artificial Analysis Intelligence Index v4.1.1 — independent lab, not a vendor — Astra scores 61.2. Claude Fable 5.1 scores 65.7, Meta's Muse Spark 1.3 62.0, and Kimi K3, the open-weight Chinese model, 57. The outside scorekeeper puts Astra near the top of a tight pack, not on top of it. Drop to the normal effort setting — what most subscribers use — and the same lab has Astra at 50 to Claude Opus 5's 51, a model costing half as much ($5/$25).
WHERE ASTRA GENUINELY LEADS: COMPUTER USE. This part isn't in dispute. On OSWorld 2.0, where a model operates a real desktop, Astra scores 72.6% against 65.7% for GPT-5.6 Sol, finishing in roughly 40 minutes where the previous model took 75. On ScreenSpot-Pro it posts 92.7% to 76.9%; on Agents' Last Exam, 59.3% to 53.6%. OpenAI president Greg Brockman called computer use "a particularly important part of what's new." That is a different product category: the difference between a tool that explains how to cancel a subscription and one that cancels it.
THREE THINGS CIRCULATING THAT ARE WRONG. First, Astra is not the strongest coding model available — Meta's Muse Spark 1.3 reportedly posts 75.4% on DeepSWE to Astra's 74.1%, the benchmark OpenAI cited. Second, the near-perfect academic scores are max-effort runs, and OpenAI part-funded FrontierMath, one of the tests it aced. Third, Humanity's Last Exam — the one major academic benchmark Astra loses, 57.2% to Fable 5.1's 65.0% — is absent from the announcement.
THE CYBERSECURITY RATING. Astra is the first model OpenAI has classified Critical for cybersecurity under its own Preparedness Framework: without safeguards it found previously unknown zero-day vulnerabilities and achieved arbitrary code execution in hardened browsers. The public version refuses proof-of-concept exploit creation and advanced offensive work; approved defensive access runs through a separate program. OpenAI also notes Astra's reasoning is harder to monitor than its predecessor's.
WHAT THIS MEANS FOR YOU: if you use a chatbot for email, research and planning, you will not feel the gap between 61 and 65 — nobody does. What's new is a model that operates software on your behalf — useful, and newly consequential, because a confident mistake now happens inside your account, not in a chat window. Practically: if you pay for ChatGPT you have Astra — try computer use on something boring and reversible first. If you pay per token, compare carefully: at standard settings Astra at $50 per million output and Claude Opus 5 at $25 are within a point of each other. And when a company says its product runs on "the world's most intelligent AI," you now know what that claim is worth.
Would you let an AI click around inside your own accounts — email, bank, calendar? Or is watching it type as far as you'd go? Tell me in the comments.
If you want AI news with the numbers actually checked, like this video, subscribe and hit the bell. Share it with someone paying for an AI tool. Patreon link below.
Keywords: GPT-6 Astra, GPT 6 review, OpenAI GPT-6, Astra vs Claude, Claude Fable 5.1, Claude Opus 5, AI benchmarks, Artificial Analysis, computer use AI, OSWorld, ChatGPT Plus, AI pricing, Kimi K3, artificial intelligence news
👇 Connect with The AI Guide:
Website: www.theaiguide.ai
YouTube: https://www.youtube.com/channel/UCamFJyTjb_kqUbNpVUVwQeg
Patreon: www.patreon.com/theaiguide
Facebook: www.facebook.com/davidtheaiguide
Instagram: www.instagram.com/theaiguide
LinkedIn: https://www.linkedin.com/company/the-ai-guide-on-youtube/
🔗 Sources:
[1] OpenAI: https://openai.com/index/gpt-6-astra/
[2] Artificial Analysis: https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra
[3] Artificial Analysis: https://artificialanalysis.ai/models/comparisons/gpt-6-astra-medium-vs-claude-opus-5
[4] Vellum: https://www.vellum.ai/blog/gpt-6-astra-benchmarks-explained
[5] Fortune: https://fortune.com/2026/09/03/openai-debuts-gpt-6-astra-computer-use-greg-brockman-says-start-of-agi/
[6] Al Jazeera: https://www.aljazeera.com/economy/2026/9/4/openai-unveils-gpt-6-astra-amid-rising-scrutiny-and-safety
#ArtificialIntelligence #OpenAI #GPT6 #AInews #TheAIGuide
bottom of page