SIGNAL.DROP
MON · WED · FRI
Drop 8.1Tuesday, Sep 29 · DevDay notes

OpenAI is building a place for agents to work.

OpenAI DevDay 2026 keynote replay · 53:37 · Sep 29
One keynote. No feed below.Open in YouTube ↗
DevDay notes, for you

01What changed

OpenAI's DevDay moved the center of gravity from a better chat reply to ongoing work. Dots are always-on agents with connected apps; ChatGPT Space gives people and agents a shared project home; the Agents API adds computer use to a managed harness. On the builder side, Codex can keep working in reusable cloud environments across devices. OpenAI also introduced GPT-6.1 Sol, claiming near-Astra performance at one-fifth Astra's standard input and output token prices. The Decisions API narrows AI judgment to finite, predefined answers, with limited preview access.

Availability is not uniform: Dots are for eligible Pro and Business Premium users, with Enterprise, Edu and Healthcare beta subject to admin enablement; Space is on Pro, Business and Enterprise. Some launches are previews or coming soon. These notes draw on OpenAI's official event recap alongside the replay; they are not a transcript or a claim to have reviewed all 53 minutes. The launch claims are not measured results in your environment.

02Why it matters to you

The paradigm shift is distribution plus durable context. An agent becomes more useful when it can remember a project, act in tools, and return to the same shared place; it also becomes harder to govern because access, state and accountability persist. For your reset-wall thesis, cheaper capable inference changes the unit economics but does not settle demand, adoption or who earns the margin. The harder question is whether persistent agents produce enough dependable work to justify the infrastructure and subscription spend.

OpenAI sells the models, workspace and runtime; the keynote is a product launch, not an independent productivity study. The visible pricing and access tiers still need a use-case test before any buying or migration decision.

03Where to apply it at work

Near-term opportunity: put the new agent-workspace pitch beside a real Now Assist or Control Tower workflow. Practical experiment: choose one bounded, recurring job and test whether the agent gets the correct final record into the system, with named access, review, exception path and rollback. Compare quality, reviewer time and cost per resolved case against the current process. A shared space and a faster model are useful only if the work lands and someone can show the receipts. ServiceNow is named among the first 32 OpenAI Marketplace partners, but the listing alone says nothing about integration depth or customer results.

Notes grounded in OpenAI's official recap, Sol model details, and computer-use guidance. Axios coverage adds an outside check; not a verification of OpenAI's outcome claims.

Say it out loudThe race is shifting from who has the smartest model to who can give an agent a job, a workspace, and a receipt when it finishes.
The running signal · through Drop 8.1

The opportunity is real. The terms decide who wins.

Eight regular picks and this DevDay addendum test the gap between promise and proof. OpenAI now packages persistent agents and shared workspaces, making the quality of the handoff, permissions and recorded outcome even more important. This is a lens to test, not a market call.

CapitalLook past the AI story to credit spreads, mortgage rates, debt terms and the shape of demand. A durable asset can have fragile financing.
AdoptionA persistent agent or benchmark score is not a business result. Check permissions, stored state, exceptions, review effort, cost and rollback before scaling.
JudgmentDon't pay extra for certainty or move with the crowd at an extreme. Buy room to learn through smaller, reversible bets.
The lineage

How the argument got here

One current video above. Open a past entry for its point, or read the full original drop below.

001The AI buildout, bull caseBrad Gerstner · Sep 13

The opening case for AI infrastructure investment: test the demand and value behind the optimism, not just the scale of the spending.

Read full drop ↓
002The credit market pushes backEd Zitron · Sep 14

The bear case: capex can race ahead of cash flow. Credit spreads and data-center debt deserve attention before equity sentiment catches up. Zitron is a polemicist; separate his argument from credit evidence.

Read full drop ↓
003Housing is more than the FedHousingWire · Sep 16

Mortgage rates answer to the bond market as well as policy rates. Follow the full financing chain rather than assuming one Fed move creates a housing buy signal.

Read full drop ↓
004The pilot-to-value gapBCG X · Sep 18

AI pilots do not prove scaled value. Put adoption, workflow change, exceptions and business outcomes on the same scorecard. BCG sells transformation work.

Read full drop ↓
005Don't pay for certaintyRichard Thaler · Sep 21

Behavioral finance's draft lesson: overpaying for a sure thing can beat neither diversification nor several measured chances. Apply that to investing and staged AI rollout.

Read full drop ↓
006The builders ask for inspectorsDario Amodei · Sep 23

Frontier pacing raises political permission as a capex input. Independent evaluators also make governance and proof of behavior part of the product, not an afterthought.

Read full drop ↓
007Underwrite the debt, not the moodHoward Marks · Sep 25

A sound technology can sit on weak financing. Marks's cycle lens turns the question from "is AI real?" to "do the terms survive a harder market?"

Read full drop ↓
008Is the score grading the right thing?Vincent Sunn Chen · Sep 28

Agent benchmarks need realistic tasks, correct final state and outcome measures. A convincing score can still miss the workflow it is meant to predict.

Read full drop ↓
8.1Give the agent a workspaceOpenAI DevDay · Sep 29 · current

Dots, Spaces and computer-use agents bring ongoing work into one vendor's surface. The test is still the correctly landed outcome, governed access and measurable cost.

Back to the running signal ↑
Full archive

Every drop, in full

Original summaries as published. Older news and numbers reflect the day of each drop, not current conditions. The past videos open outside this page, so there is still only one player here.

001Gerstner calls the doom week “hyperbolic scare tactics.”Sunday, Sep 13 - pilot
CNBC Halftime Report · Brad Gerstner, Altimeter Capital · 21 min · aired Fri Sep 11

01What happened this week

An Anthropic safety researcher quit and posted that AI labs are “gambling with our lives.” The doom discourse spiked all week. Gerstner - whose firm is invested in both OpenAI and Anthropic - went on Halftime and called it hyperbolic scare tactics hiding behind a political agenda, arguing more safety work has gone into AI than any technology in his 25 years in Silicon Valley.

02For your thesis

You track the capex reset wall - whether revenue ever catches the buildout. This fight is the other wall: whether the public lets the buildout continue at all. Political permission is an input to your 2027/28 wall. A backlash that hardens into regulation slows capex as surely as an ROI miss. The tell isn't whether the doomers are right - it's whether “gambling with our lives” starts showing up in Washington hearings.

03For the day job

Your buyers' employees are marinating in this exact discourse. Some of them now arrive pre-spooked when they hear “Now Assist.” That's why evals and governance aren't just quality gates - they're the receipts that answer this fear. The adoption pitch writes itself: we measure, we gate, we show the receipts.

Say it out loud The bull case and the doom case have become the same argument - both sides agree this thing is powerful enough to matter. They're only fighting over who's allowed to keep building.
Watch original video ↗
002The credit desks just joined the bubble fight.Monday, Sep 14
The Tech Report · Ed Zitron, Where’s Your Ed At · 44 min · posted Fri Sep 11

01What happened this weekend

Ed Zitron - the sharpest-tongued AI bear going - went on The Tech Report to argue most investors have no idea what they actually own in the AI trade. The same weekend, the adults in the room quietly conceded the shape of his argument: S&P warned hyperscaler credit quality is "gradually weakening," with capex rising faster than anticipated and financings getting "more complicated and less transparent." S&P now expects six companies - Amazon, Microsoft, Alphabet, Oracle, SpaceX, Meta - to spend $7 trillion+ on AI through 2030. And Goldman’s high-yield data-center basket is trading at 353bps, wider than June 2022 levels.

02For your thesis, and your wallet

The reset wall was always a financing story wearing an equity costume. This week the credit market said it first - spreads are pricing the wall before stocks are. Lenders are paid to be pessimistic earlier than shareholders, so that sequencing is the signal. Zitron is the sentiment; S&P and the 353bps print are the receipts. Two wealth angles: Bloomberg reports this morning that data-center debt is migrating into CMBS - the real-estate bond market - so your "what survives every scenario" asset question just got a new address. And S&P’s line that risks to the sector are risks for everyone is your index-concentration question in official language. Caveat on Zitron: polemicist, not credit analyst. Use him for what the crowd feels, the desks for what the money thinks.

03For the day job

Bubble anxiety is about to walk into your customer meetings - CFOs read Axios too. When a buyer asks whether AI spend is a bubble, the wrong answer is defending the technology. The right answer is that their spend is gated, measured, and reversible. That is literally what your evals and governance work sells. The pitch for a spooked quarter: we don’t ask you to believe, we show you the receipts.

Say it out loud The stock market is still arguing about whether AI is a bubble. The bond market already started charging for the answer.
Watch original video ↗
003Housing just became a patience trade.Wednesday, Sep 16
HousingWire · Logan Mohtashami · 23 min · posted Mon Aug 31

01What changed

The waiting game got longer. HousingWire lead analyst Logan Mohtashami’s clean frame is that mortgage rates are a bond-market story first: watch nominal growth, inflation, oil and the 10-year Treasury, not the headline Fed move. Yesterday’s data made that practical. HousingWire’s locked 30-year rate hit 7.28%, up 22 basis points in two weeks, with the 10-year near 5%. A Reuters poll published Tuesday says the housing revival will stay elusive, with rates falling only modestly and home-price growth muted through next year.

02Why it matters for your wealth

Paradigm shift: stop treating “the Fed moved” as a housing thesis. The rate buyers actually pay is the output of a system - growth, inflation, Treasury supply, risk spreads and lender capacity. Near-term opportunity: a frozen market does not make every house cheap; it makes motivated sellers, assumable debt and durable cash flow more valuable. This is not a buy signal. Mohtashami is a housing analyst, and HousingWire serves the mortgage industry. Use the framework, then underwrite the asset: price, debt, taxes, insurance and the cost to wait.

03Where to apply it at work

Practical experiment: replace single-variable stories with a small causal map. For any Now Assist rollout, write down the system before the metric - workflow volume, model quality, human adoption, exception rates and downstream outcomes - then ask what could improve the first measure while hurting the last. That is the same mistake in a different market: confusing the policy rate with the mortgage rate, or usage with value. The receipts need to show the whole chain.

Say it out loud The Fed sets one rate. The house you want is priced by an entire system.
Watch original video ↗
004The pilot is not the product.Friday, Sep 18
Omni Talk Retail · Nicolas de Bellefonds, BCG X · 14 min · posted Wed Jun 24

01What changed

The enterprise-AI story has split in two: early value is common, scaled value is not. In this sharp 14-minute interview, BCG X’s Nicolas de Bellefonds says 75% of consumer-goods companies remain stuck in pilot mode and only 18% have scaled impact. BCG’s August update widens the aperture: 82% of CEOs are more optimistic about AI ROI than a year ago, yet just 6% of companies report meaningful cost or revenue value. The gap is no longer access to models. It is the ability to change the work around them.

02Why it matters

Paradigm shift: the unit of AI transformation is the workflow, not the use case. Leaders pick a few high-value “lighthouses,” redesign them end to end, and build the shared data, talent and governance needed to repeat the move. Near-term opportunity: the scarce asset is becoming organizational follow-through - the capacity to turn a demo into a changed operating system. Caveat: BCG sells transformation work, and Omni Talk serves the retail industry. Treat the percentages as directional evidence, not neutral law. The stronger signal is that BCG’s newer research repeats the same mechanism: focus, workflow redesign and adoption separate value from activity.

03Where to apply it

Your Now Assist receipts should prove a lighthouse crossed the gap. Pick one workflow with an executive owner and a hard outcome; capture the baseline; define quality, adoption, exceptions and business value before launch; then measure the full chain after launch. A good eval says the answer was correct. A transformation receipt says the work changed, people used it, the exceptions stayed safe, and the business outcome moved. That is a much harder story for a buyer to dismiss as “another pilot.”

Say it out loud Most companies do not have an AI-model problem. They have a demo-to-operating-system problem.
Watch original video ↗
005Conviction is expensive. Options are cheap.Monday, Sep 21
Facts vs Feelings · Richard Thaler · 47 min · posted Wed Jun 17

01What the research says

Richard Thaler’s cleanest lesson is hiding in the NFL draft: the people with the strongest incentives, best data and most experience still overpay for the right to be certain. His work with Cade Massey found early picks systematically overvalued; twenty years later, teams’ ability to choose the better player has barely moved, from roughly 52% to 53%. The winning move is often to trade down, collect more shots, and admit the forecast is noisier than the price implies.

02Why it matters now

Paradigm shift: expertise does not remove bias when status rewards the bold pick. It can make the story more persuasive. After the Fed’s first hike since 2023, every market narrative now sounds certain - more hikes, no landing, buy the dip, sell duration. Thaler’s frame is better: ask what price you are paying for conviction, then prefer choices that survive being wrong. For wealth, that means diversification, staged entries and rules set before emotion arrives. Caveat: Carson Group is a wealth manager and the hosts sell advice. Thaler’s claims here are stronger where they trace to published research than where the conversation moves into current markets.

03Where to apply it

Practical experiment: build the “trade down” option into AI decisions. Instead of one big rollout justified by one forecast, run several small, measurable workflow bets with the same eval spine; fund the ones that earn evidence; stop the ones that do not. Then borrow Thaler’s retirement-design lesson: do not rely on people to make the right choice every time. Make the safer, measured path the default, with opt-outs and visible exceptions. Governance becomes choice architecture, not paperwork.

Say it out loud The smartest people in the room still overpay to feel certain. Buy more chances instead.
Watch original video ↗
006The AI labs asked for inspectors.Wednesday, Sep 23
CBS Sunday Morning · Dario Amodei · 24 min · posted Sun Sep 13

01What happened

On Sep 12, Anthropic CEO Dario Amodei published “We Must Pace the Frontier,” and within days Sam Altman, Elon Musk, Demis Hassabis and Satya Nadella backed some version of it. In this 24-minute CBS sit-down he explains it himself: the capability curve is getting steeper, “it doesn’t mean we need to panic” or shut it down, but it is a warning sign. Pacing does not mean stopping. It means every model generation gets properly tested. His first concrete step: permanent third-party evaluators inside Anthropic with badges, desks and access close to its own risk team - “like a food inspector.” Chip stocks fell (SMH down 5.5% for the week through Sep 15). Trump called the safety worries a hoax, and Jensen Huang argued a slowdown helps China.

02What it means for your thesis

Paradigm shift: political permission just showed up as a visible input to the capex curve, and this time the labs raised it, not regulators. Near-term: the market priced a spending pullback that nobody actually proposed, and no lab has named a model it will delay. The sharper risk is a mismatch. Labs can change training plans in weeks; land, transformers and substations are committed for years. Demand may move from giant training campuses to spread-out inference instead of vanishing. That tightens the reset-wall question: not “does demand disappear,” but “does it arrive in the shape the debt was underwritten for.” Powered land in tight markets still looks scarce (CBRE: 1.4% vacancy, 80% of new builds preleased); single-purpose training sites carry more of the risk. Caveat: Anthropic helps set the evaluator standard it is asking everyone to adopt, and VanEck reports it is expected to start IPO marketing next month. VanEck sells chip ETFs. Huang sells the chips. Everyone in this debate is talking their book.

03What it means for your work

The frontier just adopted your pitch: governance is the receipts. An embedded evaluator is continuous, independent evals plus a real incident path, which is exactly what Now Assist customers will start asking for once the headline filters into procurement. Practical experiment: for one live workflow, build the “inspector’s desk” - the eval set, logs, incident path and rollback an outsider would need to verify it behaves. Build that before you scale it. Near-term opportunity: if the frontier slows, the demand that holds up is enterprise inference judged on measured value, not next-model hype. Gartner’s advice fits the pitch: staged rollouts, flexible capacity, regular reassessment.

Say it out loud The AI labs just asked for food inspectors. When the builders want inspectors, the receipts become the product.
Watch original video ↗
007The worst loans are made in the best times.Friday, Sep 25
Howard Marks on TBPN · 36 min · recorded Feb 26; published May 26

01What happened

In this February interview, Oaktree's Howard Marks gives you a clean cycle model: underlying value can move steadily while prices swing above and below it as psychology changes. Oaktree mostly invests in credit, not public stocks. Its posture flips with the crowd: cautious when lenders act as if risk has vanished, more willing to lend or buy distressed debt when fear finally pays you to take risk. His line is blunt: “The worst of loans are made in the best of times.”

This is an older recording, not a reaction to this week's market. The fresh reason to revisit it: mortgage rates jumped on Sep 24 as bond yields rose. CNBC cited Mortgage News Daily's 7.45% daily reading; Freddie Mac's weekly survey runs on a different, lagged window. The mechanism is the point, not a single rate print.

02What it means for you

For your reset-wall thesis, separate three questions: is the technology useful, is the asset durable, and was its debt priced for the hard case? Those can have different answers. If the growth curve is real but cash flows arrive later or in a different shape, a sound asset can still sit on fragile financing. With powered land or data-center REITs, examine tenant credit, contracted demand, debt maturities, power delivery and refinancing terms before calling something “scarce and enduring.”

Marks is a credit manager whose firm benefits from distressed opportunities. His cycle lens is useful, but it is not a forecast that a crash is due or a reason to buy a particular security.

03What it means for your work

The same split helps with Now Assist: capability is not customer value, and a strong demo is not durable adoption. Marks says AI can gather evidence and frame hypotheses, while humans still need to test unfamiliar cases. Practical experiment: take one live workflow and log the baseline, time saved, rework, cost per resolved case, failure cases and rollback owner. Ask whether those receipts would still make the deal compelling with a slower model cycle or a tighter budget. That is a stronger adoption story than enthusiasm alone.

Say it out loudThe tech can be real and the financing can still be wrong. In a boom, underwrite the debt, not the mood.
Watch original video ↗
008The score might be grading the wrong thing.Monday, Sep 28
Snorkel AI · Vincent Sunn Chen · 13 min · published Aug 3

01What changed

In this short August talk, Snorkel's Vincent Sunn Chen asks what makes an agent benchmark last: realistic tasks, long enough work horizons, and judging the quality of the result, not just a neat answer. His example, Senior SWE-Bench, uses underspecified engineering jobs and checks behavior plus code quality. Newer research sharpens the question. Stanford reported on Sep 25 that tests claiming to measure the same AI quality can disagree; an Accenture ERP study found agents could reach Save yet leave the business record wrong. Those later studies are context for today's pick, not claims made in the video.

02Why it matters to you

Paradigm shift: the measuring stick is part of the product. A leaderboard score is not the same as a dependable workflow or a return on investment. The money angle is the same as your reset-wall question: if funding assumes agents replace work at scale, test whether the claimed productivity survives realistic tasks, exceptions, review time and rework. Near-term opportunity: teams that can prove real outcomes have a better story than teams selling benchmark points.

Snorkel sells AI data and evaluation tools, and its benchmark is a coding benchmark, not a direct score for Now Assist. Accenture's ERP results are one research setting, not a general failure rate.

03Where to apply it at work

Practical experiment: take one Now Assist workflow and write down the intended final state before testing. Include a normal case, an underspecified request and an exception. After each run, inspect the stored record and downstream action, not just the click path or a model judge's score. Have a reviewer grade usefulness, rework and catch rate, then compare cost per correctly resolved case with the current process. If the score says “pass” but the record or outcome is wrong, fix the measuring stick before scaling the agent.

Say it out loudIf the agent clicked Save but saved the wrong thing, the benchmark passed the wrong exam.
Watch original video ↗
8.1OpenAI is building a place for agents to workDevDay notes · Sep 29 · current
OpenAI DevDay 2026 keynote replay · 53:37 · Sep 29

01What changed

OpenAI's DevDay moved the center of gravity from a better chat reply to ongoing work. Dots are always-on agents with connected apps; ChatGPT Space gives people and agents a shared project home; the Agents API adds computer use to a managed harness. On the builder side, Codex can keep working in reusable cloud environments across devices. OpenAI also introduced GPT-6.1 Sol, claiming near-Astra performance at one-fifth Astra's standard input and output token prices. The Decisions API narrows AI judgment to finite, predefined answers, with limited preview access.

Availability is not uniform: Dots are for eligible Pro and Business Premium users, with Enterprise, Edu and Healthcare beta subject to admin enablement; Space is on Pro, Business and Enterprise. Some launches are previews or coming soon. These notes draw on OpenAI's official event recap alongside the replay; they are not a transcript or a claim to have reviewed all 53 minutes. The launch claims are not measured results in your environment.

02Why it matters to you

The paradigm shift is distribution plus durable context. An agent becomes more useful when it can remember a project, act in tools, and return to the same shared place; it also becomes harder to govern because access, state and accountability persist. For your reset-wall thesis, cheaper capable inference changes the unit economics but does not settle demand, adoption or who earns the margin. The harder question is whether persistent agents produce enough dependable work to justify the infrastructure and subscription spend.

OpenAI sells the models, workspace and runtime; the keynote is a product launch, not an independent productivity study. The visible pricing and access tiers still need a use-case test before any buying or migration decision.

03Where to apply it at work

Near-term opportunity: put the new agent-workspace pitch beside a real Now Assist or Control Tower workflow. Practical experiment: choose one bounded, recurring job and test whether the agent gets the correct final record into the system, with named access, review, exception path and rollback. Compare quality, reviewer time and cost per resolved case against the current process. A shared space and a faster model are useful only if the work lands and someone can show the receipts. ServiceNow is named among the first 32 OpenAI Marketplace partners, but the listing alone says nothing about integration depth or customer results.

Say it out loudThe race is shifting from who has the smartest model to who can give an agent a job, a workspace, and a receipt when it finishes.

These notes draw on OpenAI's official recap; the entire keynote was not reviewed as a transcript. Announced capabilities are not measured results. OpenAI sells the products.

Watch keynote replay ↗
Read official recap ↗