OpenAI's DevDay moved the center of gravity from a better chat reply to ongoing work. Dots are always-on agents with connected apps; ChatGPT Space gives people and agents a shared project home; the Agents API adds computer use to a managed harness. On the builder side, Codex can keep working in reusable cloud environments across devices. OpenAI also introduced GPT-6.1 Sol, claiming near-Astra performance at one-fifth Astra's standard input and output token prices. The Decisions API narrows AI judgment to finite, predefined answers, with limited preview access.
Availability is not uniform: Dots are for eligible Pro and Business Premium users, with Enterprise, Edu and Healthcare beta subject to admin enablement; Space is on Pro, Business and Enterprise. Some launches are previews or coming soon. These notes draw on OpenAI's official event recap alongside the replay; they are not a transcript or a claim to have reviewed all 53 minutes. The launch claims are not measured results in your environment.
The paradigm shift is distribution plus durable context. An agent becomes more useful when it can remember a project, act in tools, and return to the same shared place; it also becomes harder to govern because access, state and accountability persist. For your reset-wall thesis, cheaper capable inference changes the unit economics but does not settle demand, adoption or who earns the margin. The harder question is whether persistent agents produce enough dependable work to justify the infrastructure and subscription spend.
OpenAI sells the models, workspace and runtime; the keynote is a product launch, not an independent productivity study. The visible pricing and access tiers still need a use-case test before any buying or migration decision.
Near-term opportunity: put the new agent-workspace pitch beside a real Now Assist or Control Tower workflow. Practical experiment: choose one bounded, recurring job and test whether the agent gets the correct final record into the system, with named access, review, exception path and rollback. Compare quality, reviewer time and cost per resolved case against the current process. A shared space and a faster model are useful only if the work lands and someone can show the receipts. ServiceNow is named among the first 32 OpenAI Marketplace partners, but the listing alone says nothing about integration depth or customer results.
Eight regular picks and this DevDay addendum test the gap between promise and proof. OpenAI now packages persistent agents and shared workspaces, making the quality of the handoff, permissions and recorded outcome even more important. This is a lens to test, not a market call.
One current video above. Open a past entry for its point, or read the full original drop below.
The opening case for AI infrastructure investment: test the demand and value behind the optimism, not just the scale of the spending.
Read full drop ↓The bear case: capex can race ahead of cash flow. Credit spreads and data-center debt deserve attention before equity sentiment catches up. Zitron is a polemicist; separate his argument from credit evidence.
Read full drop ↓Mortgage rates answer to the bond market as well as policy rates. Follow the full financing chain rather than assuming one Fed move creates a housing buy signal.
Read full drop ↓AI pilots do not prove scaled value. Put adoption, workflow change, exceptions and business outcomes on the same scorecard. BCG sells transformation work.
Read full drop ↓Behavioral finance's draft lesson: overpaying for a sure thing can beat neither diversification nor several measured chances. Apply that to investing and staged AI rollout.
Read full drop ↓Frontier pacing raises political permission as a capex input. Independent evaluators also make governance and proof of behavior part of the product, not an afterthought.
Read full drop ↓A sound technology can sit on weak financing. Marks's cycle lens turns the question from "is AI real?" to "do the terms survive a harder market?"
Read full drop ↓Agent benchmarks need realistic tasks, correct final state and outcome measures. A convincing score can still miss the workflow it is meant to predict.
Read full drop ↓Dots, Spaces and computer-use agents bring ongoing work into one vendor's surface. The test is still the correctly landed outcome, governed access and measurable cost.
Back to the running signal ↑Original summaries as published. Older news and numbers reflect the day of each drop, not current conditions. The past videos open outside this page, so there is still only one player here.
An Anthropic safety researcher quit and posted that AI labs are “gambling with our lives.” The doom discourse spiked all week. Gerstner - whose firm is invested in both OpenAI and Anthropic - went on Halftime and called it hyperbolic scare tactics hiding behind a political agenda, arguing more safety work has gone into AI than any technology in his 25 years in Silicon Valley.
You track the capex reset wall - whether revenue ever catches the buildout. This fight is the other wall: whether the public lets the buildout continue at all. Political permission is an input to your 2027/28 wall. A backlash that hardens into regulation slows capex as surely as an ROI miss. The tell isn't whether the doomers are right - it's whether “gambling with our lives” starts showing up in Washington hearings.
Your buyers' employees are marinating in this exact discourse. Some of them now arrive pre-spooked when they hear “Now Assist.” That's why evals and governance aren't just quality gates - they're the receipts that answer this fear. The adoption pitch writes itself: we measure, we gate, we show the receipts.
Ed Zitron - the sharpest-tongued AI bear going - went on The Tech Report to argue most investors have no idea what they actually own in the AI trade. The same weekend, the adults in the room quietly conceded the shape of his argument: S&P warned hyperscaler credit quality is "gradually weakening," with capex rising faster than anticipated and financings getting "more complicated and less transparent." S&P now expects six companies - Amazon, Microsoft, Alphabet, Oracle, SpaceX, Meta - to spend $7 trillion+ on AI through 2030. And Goldman’s high-yield data-center basket is trading at 353bps, wider than June 2022 levels.
The reset wall was always a financing story wearing an equity costume. This week the credit market said it first - spreads are pricing the wall before stocks are. Lenders are paid to be pessimistic earlier than shareholders, so that sequencing is the signal. Zitron is the sentiment; S&P and the 353bps print are the receipts. Two wealth angles: Bloomberg reports this morning that data-center debt is migrating into CMBS - the real-estate bond market - so your "what survives every scenario" asset question just got a new address. And S&P’s line that risks to the sector are risks for everyone is your index-concentration question in official language. Caveat on Zitron: polemicist, not credit analyst. Use him for what the crowd feels, the desks for what the money thinks.
Bubble anxiety is about to walk into your customer meetings - CFOs read Axios too. When a buyer asks whether AI spend is a bubble, the wrong answer is defending the technology. The right answer is that their spend is gated, measured, and reversible. That is literally what your evals and governance work sells. The pitch for a spooked quarter: we don’t ask you to believe, we show you the receipts.
The waiting game got longer. HousingWire lead analyst Logan Mohtashami’s clean frame is that mortgage rates are a bond-market story first: watch nominal growth, inflation, oil and the 10-year Treasury, not the headline Fed move. Yesterday’s data made that practical. HousingWire’s locked 30-year rate hit 7.28%, up 22 basis points in two weeks, with the 10-year near 5%. A Reuters poll published Tuesday says the housing revival will stay elusive, with rates falling only modestly and home-price growth muted through next year.
Paradigm shift: stop treating “the Fed moved” as a housing thesis. The rate buyers actually pay is the output of a system - growth, inflation, Treasury supply, risk spreads and lender capacity. Near-term opportunity: a frozen market does not make every house cheap; it makes motivated sellers, assumable debt and durable cash flow more valuable. This is not a buy signal. Mohtashami is a housing analyst, and HousingWire serves the mortgage industry. Use the framework, then underwrite the asset: price, debt, taxes, insurance and the cost to wait.
Practical experiment: replace single-variable stories with a small causal map. For any Now Assist rollout, write down the system before the metric - workflow volume, model quality, human adoption, exception rates and downstream outcomes - then ask what could improve the first measure while hurting the last. That is the same mistake in a different market: confusing the policy rate with the mortgage rate, or usage with value. The receipts need to show the whole chain.
The enterprise-AI story has split in two: early value is common, scaled value is not. In this sharp 14-minute interview, BCG X’s Nicolas de Bellefonds says 75% of consumer-goods companies remain stuck in pilot mode and only 18% have scaled impact. BCG’s August update widens the aperture: 82% of CEOs are more optimistic about AI ROI than a year ago, yet just 6% of companies report meaningful cost or revenue value. The gap is no longer access to models. It is the ability to change the work around them.
Paradigm shift: the unit of AI transformation is the workflow, not the use case. Leaders pick a few high-value “lighthouses,” redesign them end to end, and build the shared data, talent and governance needed to repeat the move. Near-term opportunity: the scarce asset is becoming organizational follow-through - the capacity to turn a demo into a changed operating system. Caveat: BCG sells transformation work, and Omni Talk serves the retail industry. Treat the percentages as directional evidence, not neutral law. The stronger signal is that BCG’s newer research repeats the same mechanism: focus, workflow redesign and adoption separate value from activity.
Your Now Assist receipts should prove a lighthouse crossed the gap. Pick one workflow with an executive owner and a hard outcome; capture the baseline; define quality, adoption, exceptions and business value before launch; then measure the full chain after launch. A good eval says the answer was correct. A transformation receipt says the work changed, people used it, the exceptions stayed safe, and the business outcome moved. That is a much harder story for a buyer to dismiss as “another pilot.”
Richard Thaler’s cleanest lesson is hiding in the NFL draft: the people with the strongest incentives, best data and most experience still overpay for the right to be certain. His work with Cade Massey found early picks systematically overvalued; twenty years later, teams’ ability to choose the better player has barely moved, from roughly 52% to 53%. The winning move is often to trade down, collect more shots, and admit the forecast is noisier than the price implies.
Paradigm shift: expertise does not remove bias when status rewards the bold pick. It can make the story more persuasive. After the Fed’s first hike since 2023, every market narrative now sounds certain - more hikes, no landing, buy the dip, sell duration. Thaler’s frame is better: ask what price you are paying for conviction, then prefer choices that survive being wrong. For wealth, that means diversification, staged entries and rules set before emotion arrives. Caveat: Carson Group is a wealth manager and the hosts sell advice. Thaler’s claims here are stronger where they trace to published research than where the conversation moves into current markets.
Practical experiment: build the “trade down” option into AI decisions. Instead of one big rollout justified by one forecast, run several small, measurable workflow bets with the same eval spine; fund the ones that earn evidence; stop the ones that do not. Then borrow Thaler’s retirement-design lesson: do not rely on people to make the right choice every time. Make the safer, measured path the default, with opt-outs and visible exceptions. Governance becomes choice architecture, not paperwork.
On Sep 12, Anthropic CEO Dario Amodei published “We Must Pace the Frontier,” and within days Sam Altman, Elon Musk, Demis Hassabis and Satya Nadella backed some version of it. In this 24-minute CBS sit-down he explains it himself: the capability curve is getting steeper, “it doesn’t mean we need to panic” or shut it down, but it is a warning sign. Pacing does not mean stopping. It means every model generation gets properly tested. His first concrete step: permanent third-party evaluators inside Anthropic with badges, desks and access close to its own risk team - “like a food inspector.” Chip stocks fell (SMH down 5.5% for the week through Sep 15). Trump called the safety worries a hoax, and Jensen Huang argued a slowdown helps China.
Paradigm shift: political permission just showed up as a visible input to the capex curve, and this time the labs raised it, not regulators. Near-term: the market priced a spending pullback that nobody actually proposed, and no lab has named a model it will delay. The sharper risk is a mismatch. Labs can change training plans in weeks; land, transformers and substations are committed for years. Demand may move from giant training campuses to spread-out inference instead of vanishing. That tightens the reset-wall question: not “does demand disappear,” but “does it arrive in the shape the debt was underwritten for.” Powered land in tight markets still looks scarce (CBRE: 1.4% vacancy, 80% of new builds preleased); single-purpose training sites carry more of the risk. Caveat: Anthropic helps set the evaluator standard it is asking everyone to adopt, and VanEck reports it is expected to start IPO marketing next month. VanEck sells chip ETFs. Huang sells the chips. Everyone in this debate is talking their book.
The frontier just adopted your pitch: governance is the receipts. An embedded evaluator is continuous, independent evals plus a real incident path, which is exactly what Now Assist customers will start asking for once the headline filters into procurement. Practical experiment: for one live workflow, build the “inspector’s desk” - the eval set, logs, incident path and rollback an outsider would need to verify it behaves. Build that before you scale it. Near-term opportunity: if the frontier slows, the demand that holds up is enterprise inference judged on measured value, not next-model hype. Gartner’s advice fits the pitch: staged rollouts, flexible capacity, regular reassessment.
In this February interview, Oaktree's Howard Marks gives you a clean cycle model: underlying value can move steadily while prices swing above and below it as psychology changes. Oaktree mostly invests in credit, not public stocks. Its posture flips with the crowd: cautious when lenders act as if risk has vanished, more willing to lend or buy distressed debt when fear finally pays you to take risk. His line is blunt: “The worst of loans are made in the best of times.”
This is an older recording, not a reaction to this week's market. The fresh reason to revisit it: mortgage rates jumped on Sep 24 as bond yields rose. CNBC cited Mortgage News Daily's 7.45% daily reading; Freddie Mac's weekly survey runs on a different, lagged window. The mechanism is the point, not a single rate print.
For your reset-wall thesis, separate three questions: is the technology useful, is the asset durable, and was its debt priced for the hard case? Those can have different answers. If the growth curve is real but cash flows arrive later or in a different shape, a sound asset can still sit on fragile financing. With powered land or data-center REITs, examine tenant credit, contracted demand, debt maturities, power delivery and refinancing terms before calling something “scarce and enduring.”
Marks is a credit manager whose firm benefits from distressed opportunities. His cycle lens is useful, but it is not a forecast that a crash is due or a reason to buy a particular security.
The same split helps with Now Assist: capability is not customer value, and a strong demo is not durable adoption. Marks says AI can gather evidence and frame hypotheses, while humans still need to test unfamiliar cases. Practical experiment: take one live workflow and log the baseline, time saved, rework, cost per resolved case, failure cases and rollback owner. Ask whether those receipts would still make the deal compelling with a slower model cycle or a tighter budget. That is a stronger adoption story than enthusiasm alone.
In this short August talk, Snorkel's Vincent Sunn Chen asks what makes an agent benchmark last: realistic tasks, long enough work horizons, and judging the quality of the result, not just a neat answer. His example, Senior SWE-Bench, uses underspecified engineering jobs and checks behavior plus code quality. Newer research sharpens the question. Stanford reported on Sep 25 that tests claiming to measure the same AI quality can disagree; an Accenture ERP study found agents could reach Save yet leave the business record wrong. Those later studies are context for today's pick, not claims made in the video.
Paradigm shift: the measuring stick is part of the product. A leaderboard score is not the same as a dependable workflow or a return on investment. The money angle is the same as your reset-wall question: if funding assumes agents replace work at scale, test whether the claimed productivity survives realistic tasks, exceptions, review time and rework. Near-term opportunity: teams that can prove real outcomes have a better story than teams selling benchmark points.
Snorkel sells AI data and evaluation tools, and its benchmark is a coding benchmark, not a direct score for Now Assist. Accenture's ERP results are one research setting, not a general failure rate.
Practical experiment: take one Now Assist workflow and write down the intended final state before testing. Include a normal case, an underspecified request and an exception. After each run, inspect the stored record and downstream action, not just the click path or a model judge's score. Have a reviewer grade usefulness, rework and catch rate, then compare cost per correctly resolved case with the current process. If the score says “pass” but the record or outcome is wrong, fix the measuring stick before scaling the agent.
OpenAI's DevDay moved the center of gravity from a better chat reply to ongoing work. Dots are always-on agents with connected apps; ChatGPT Space gives people and agents a shared project home; the Agents API adds computer use to a managed harness. On the builder side, Codex can keep working in reusable cloud environments across devices. OpenAI also introduced GPT-6.1 Sol, claiming near-Astra performance at one-fifth Astra's standard input and output token prices. The Decisions API narrows AI judgment to finite, predefined answers, with limited preview access.
Availability is not uniform: Dots are for eligible Pro and Business Premium users, with Enterprise, Edu and Healthcare beta subject to admin enablement; Space is on Pro, Business and Enterprise. Some launches are previews or coming soon. These notes draw on OpenAI's official event recap alongside the replay; they are not a transcript or a claim to have reviewed all 53 minutes. The launch claims are not measured results in your environment.
The paradigm shift is distribution plus durable context. An agent becomes more useful when it can remember a project, act in tools, and return to the same shared place; it also becomes harder to govern because access, state and accountability persist. For your reset-wall thesis, cheaper capable inference changes the unit economics but does not settle demand, adoption or who earns the margin. The harder question is whether persistent agents produce enough dependable work to justify the infrastructure and subscription spend.
OpenAI sells the models, workspace and runtime; the keynote is a product launch, not an independent productivity study. The visible pricing and access tiers still need a use-case test before any buying or migration decision.
Near-term opportunity: put the new agent-workspace pitch beside a real Now Assist or Control Tower workflow. Practical experiment: choose one bounded, recurring job and test whether the agent gets the correct final record into the system, with named access, review, exception path and rollback. Compare quality, reviewer time and cost per resolved case against the current process. A shared space and a faster model are useful only if the work lands and someone can show the receipts. ServiceNow is named among the first 32 OpenAI Marketplace partners, but the listing alone says nothing about integration depth or customer results.
These notes draw on OpenAI's official recap; the entire keynote was not reviewed as a transcript. Announced capabilities are not measured results. OpenAI sells the products.
Watch keynote replay ↗