John Cochrane gets AI-pilled
by a refined tool
I recently tried refine, an AI tool for refining academic articles, developed by Yann Calvó López and Ben Golub. I sent it the current draft of my booklet on inflation, to see what it can offer. …
The results are stunning. The comments it offered were on the par of the best comments I’ve received on a paper in my entire academic career. And more concise and organized than the best ones. They aren’t perfect, but the kind of analysis the program is able to do is past the point where technology looks like magic. I don’t know how you get here from “predict the next word.”
I am curious to know how the economists’ domain expertise contributed to the results. I asked Claude Opus 4.6 ($20 a month version) to write a referee’s report on the booklet. It said that it was only able to fetch the first four chapters, because the pdf was too long. Still, Claude wrote this:
Cochrane is refreshingly honest that the government debt valuation equation is, like the stock-price present-value identity, essentially untestable in isolation. He warns readers not to waste time constructing independent measures of the present value of surpluses. But this candor creates a deep tension with the rest of the paper, which claims that FTPL explains episodes better than alternatives. If the central variable — expected future surpluses — is unobservable, and if any observed path of data is consistent with both New-Keynesian equilibrium selection and FTPL (the observational equivalence theorem stated in Section 4.5), then on what empirical grounds does one actually prefer FTPL? The paper's answer is essentially "plausibility of narratives." That is a legitimate form of argument, but the reader deserves a more sustained and self-critical discussion of how one adjudicates narrative plausibility without sliding into confirmation bias. The analogy with asset pricing is apt but cuts both ways: the present-value identity for stocks has been extremely hard to use for discriminating among competing models of risk premia, and decades of debate remain unresolved. One might worry the same fate awaits FTPL.
This comment seems to me to be on par with what Refine wrote on this issue. And Claude writes this:
the specificity of the narrative (inflation broke out in February 2021 because of the American Rescue Plan signal; inflation peaked in mid-2022 because of expected midterm results) does a lot of load-bearing work. Could one tell an equally plausible FTPL story with different timing? If so, the theory accommodates too much. If not, the timing predictions should be testable against, say, breakeven inflation rates, CDS spreads on sovereign debt, or fiscal news events, and the paper would benefit from at least gesturing toward such evidence.
This sounds like a high-caliber, diligent referee. You certainly cannot dismiss it as slo slop.
So for me, the question of how much value added López and Golub provide is not conclusively answered. This is a big issue in the world of AI start-ups. If domain expertise adds a lot of value, then VC’s should be backing start-ups that adapt frontier models to particular domains. But there is a contrary view that these are bad investments, because the frontier models are going to be able to master particular domains without needing significant training or modifications.
Cochrane later writes,
I also tried Claude to update some graphs. My prompt was just “write a matlab program that fetches data series xyz from Fred using the API, and make a graph that..” with pretty detailed description of the graph. It ran right out of the box, even doing a decent job of “put text labels on the graph in a way that doesn’t conflict with the plotted time series.” Claude did not do a good job of finding which Fred data series would work, but that was a small task. And it produced code using a lot of commands I don’t recognize. Making sure programs do what you think they do will be a new challenge. It went on and did things I didn’t ask for, like offer summary statistics! Still, an hour job took 5 minutes.
This is all old news to most of my colleagues, who are integrating AI into workflows with great speed. But if you’re not using these tools, the time to start is now.
substacks referenced above: @



My sense is that the emergent capabilities of the frontier models will exceed efforts to specifically tune models for particular domains. The area I know the most about is the law, commercial/transactional and regulatory law in particular, and the frontier models are extremely competent in that field, and improving rapidly. My impression is that using one of the ‘wrappers’ customized for attorneys just locks you into a potentially second-tier or lagging LLM.
This is often a feature of technology. The scale of generally applicable technologies drives ‘what’s best’, as opposed to domain specialization. Every now and then there’s a niche player who ekes out a living like Etsy or GoPro. But mostly it’s Amazon and Apple who provide the best solutions.
(The history of Silicon Valley is littered with such examples – I always loved the fact that the computer history museum is located in the former headquarters of Silicon Graphics – literally now itself part of computer history. Silicon Graphics specialized in high-performance graphics hardware, which ultimately couldn’t compete against the scale of the video cards that went into general purpose PCs – brought to you, in part, by Nvidia.)
Separate but related: To my mind the economics is also different from how you describe it. Frontier labs will not throw double digit dollars of compute at a single query.
So the question for users is: conditional on spending over $20 for essentially a single query, do you want to do with a tool that is customized to root out every type of error in technical documents, or a generalist tool whose reputation doesn't rest on that capability.... It's not an obvious question, but it's a different way of looking at it.