The Wrapper Model
Have academic economists fallen for a scam?
Tyler Cowen has been tracking the success of “Refine,” which describes itself as
an AI-based tool that identifies errors in reasoning, calculation, and references across parts of a paper.
For authors and journals alike, it’s a rigorous check that provides peace of mind before work goes out into the world in its final form.
It might be a scam.
Can the tool do what it says it does? I would bet yes.
Does this provide value compared to not using any tool? I would bet yes.
Can it provide services more cost-effectively than a human? I would bet yes.
Can the tool provide significantly better results than just feeding your paper to the latest version of Claude or ChatGPT and asking it to “identify errors in reasoning, calculation, and references?” I would bet no.
The term of art for “Refine” is a wrapper model. That is, it wraps around one or more of the frontier models.
I have not played with “Refine,” because trying to publish a journal article is not my jam. But for other projects, such as my lectures on 21st century American history, I am finding it more satisfying to go directly to Claude rather than use a wrapper. I imagine that if I were interested in publishing a journal article, I could go directly to Claude and receive feedback that is at least as useful as what I could get from using a wrapper.
I know many domains in which the wrapper model is generating revenue. My guess is that the decision-makers in those domains are not AI natives. They are the equivalent of the executives of the 1990s who had their secretaries print out their emails.
The wrapper modelers had better take their profits while they can. Once AI natives rise to positions of authority, they will discard the wrapper models in favor of going directly to the frontier models.


Just to clarify, Refine has released a benchmark and outperforms 'out-of-the-box' models and open-source scaffolds by a wide margin. Creating a scaffold that combines multiple frontier models and produces excellent results is non-trivial.
https://www.refine.ink/blog/refine-ai-reviewer-benchmark
What was wrong about printing out your emails?