My guess is that pricing of top performance by big models will be such that specialized wrappers, which help minimize the cost of achieving the top performance in various legal, medical, engineering, math, programming fields will allow such specialized wrapper/ front-ends to demonstrate value added -- usually more value added than middle management.
What To Do? The key management Q. How best to do it? The engineering, workers key Q.
The middle management is to protect the decision makers from too many questions about "what", and especially to monitor the front line managers with a series of summary reports on progress towards the top decided What To Do.
Every current department of each org might be a good candidate for a very specific company - product wrapper to compare various Top ai models.
It's unlikely the best ai model will be the best in every field, all the time. For most work, the "best" answer is hardly worth more than the next best 100 answers. The best value answer is some combination of quality & price.
Quality, like in cars or anything else, has many objective features but always subjective valuations of each feature. Now I'm thinking of the classic car joke, where two guys are arguing over which car is better, based on speed, turning radius, etc. They ask a female mutual friend. She quickly answers: the Red one.
Creating a scaffold combining multiple top models to get great results seems very likely to remain valuable; and to allow a decision maker to use that ai rather than an assistant to do the comparison.
I worked for people like this: "They are the equivalent of the executives of the 1990s who had their secretaries print out their emails." A compelling analogy for folks of a certain age.
For a particular academic field, having a particular wrapper dedicated to that field gives academics collectively "someone to talk to" in terms of pointing out things that need to be fixed or, well, refined, according to the particular needs and interests that are the highest status in that field. Trying to "petition" changes in something like Claude feels like trying to petition the federal government, it's too big and too wide in scope for your complaints or suggestions to matter. But a field-specialized tool is more like a local official with whom you could actually set up a meeting. The company behind it is focusing on the opinions and usage of a particular narrow community and thus has a kind of natural closer relationship to its membership and an interest interest in paying attention to its comments and concerns and making appropriate adjustments. Claude and the others and the big companies behind them are just too big to "too inpersonal" to care about those concerns, requests, interests, and details for some tiny subgroup of its giant and diverse population of users. I can especially see the value of field-dedicated wrapper products and companies in the law. Indeed, that was part of the history of Westlaw and Lexis when the big search engines came on the scene and everybody predicted that was the end of the specialist firms. It didn't work out that way because generic search that is the 95% adequate solution for a million other users is (for various reasons I'll spare you) not a good approach to legal research, the knowledge of which is organized in a distinct kind of intellectual structure. Notice that even after several years of knowing about big problems, generic AI tools are still really bad at legal research and writing, and the legally specialist AI tools, tweaked to the particular needs of the legal profession, are (at least currently) considred clearly superior (albeit a lot more expensive, which is why lawyers try to save money and use the generic tools and are still often getting themselves into big trouble.
Using AI well is all about workflows and iteratively improved, quantitatively evaluated prompts. There is surely room for someone to come up with something much better than just feeding it straight into ChatGPT.
Whether or not Refine actually does that, I cannot say. But there is an opportunity for such a thing.
Reviewing is one of the easier, most beneficial ways that most people can benefit by using more AI.
It is great for programming, which is one of the more hyped capabilities. Even if you don't use it for generating code, having it review your patches will often give you good next steps that you would not have discovered of on your own.
It is also great for documents, though, including any form of report, email, or slide show. Similar to programming reviews, it will not just look for low-level issues such as misspellings and bad grammar, but also issues of tone and of pacing.
If you make any form of digital artifact, then having an extra reviewer tends to be valuable. If it gives you something you act on even half the time, then it's helping you produce better output, at a very low cost.
Getting back to Refine, I share the skepticism that a special-purpose tool is likely to beat direct usage of ChatGPT or Claude. The general-purpose tools learn the current best practices for basically all specialized domains. At most, you may want to add t othe prompt something like, "this is a paper intended for academic publication at (journal name)", so that it activates the relevant specialized info that it should have.
Granted, academic publication may start to really fray around the edges if the papers are generated by AI and also reviewed by them. The world is changing.
There is a popular wrapper model that uses the major LLMs for education applications called Magic School. I think much of its appeal is that it lists specific uses (draft an email to parents, design a quiz, scaffold an assignment by reading level, etc.) that AI-illiterate and uncreative teachers wouldn't think of on their own. Giving each narrow usage its own designated tool makes inputting easier for people as well.
Once AI use across education and other fields is compulsory and ubiquitous, people will probably have developed the proficiency to interact with the models themselves without the training wheels that these platforms provide.
Just to clarify, Refine has released a benchmark and outperforms 'out-of-the-box' models and open-source scaffolds by a wide margin. Creating a scaffold that combines multiple frontier models and produces excellent results is non-trivial.
I think you should make it clear that you are an advisor to Refine. Your remarks did not seem to me to be disinterested so I was not surprised to find you on https://www.refine.ink/team. I don't have a "dog in this fight" but I think you should disclose if you do.
Just to clarify, I asked Claude about this paper. (https://claude.ai/share/e62ea757-35c5-4cf8-81fd-744ab703cd83). One quote: "Their scoring procedure defines a good review as one containing many verifiable, paper-grounded, atomic concerns, and explicitly "gives less credit to broad literature positioning, fit, taste, or claims requiring outside facts." But that is precisely what Refine is built to produce. So the benchmark isn't neutrally measuring "review quality" — it's measuring "closeness to Refine's theory of what a review should be." A human referee's most valuable contribution is often exactly the things this rubric discounts: judgment about whether the paper matters, whether the framing is right, how it sits in the literature. A system optimized to emit lots of granular anchored concerns will win a contest scored on lots of granular anchored concerns almost by construction.
I understand the skepticism, and there is much to improve on. That said, Refine is being used exactly for this purpose (post-acceptance but pre-publication review) by the AEA / Econometric Society. It isn't meant to replace judgement about contribution / framing.
In terms of review philosophy and the benchmark, the competing models are given a prompt that also asks for verifiable and atomic concerns. So the benchmark is not unfairly comparing review systems with differing objectives.
Writers at the post-graduate level are probably conscientious enough to use tools like this productively, recognizing the errors the AI corrects and adapting over time. Undergraduates would input their sloppy passive-voice written papers and have the AI rewrite everything without reading any feedback. This is already happening, of course.
Platforms like Refine should add a student mode, and schools should require mandatory writing seminars where students feed their papers and essays into it. An instructor would then review the results and work with students individually based on their weaknesses in writing and reasoning. It's unlikely to happen in any case because modern American universities do not engage in productive pedagogy.
Refine sounds a bit like illusion of control on the part of the Econ profession. “We have vetted this tool, understand it, and give you permission to use it”.
When in reality everyone is using AI already in their work and quite aggressively. No one is waiting around for the installed, err… elected AEA president to tell them what tools to use.
I know it’s not the focus of this post, but as someone who retired from the publishing race in economics a long time ago, I wonder: can any of these models rank the academic economics articles on the index of highest in elegance, but lowest in relevance? :)
After reading the comments, I wonder if a dedicated peer-review referee wrapper wouldn't be an ideal product even if humans are nominally uncompensated for performing that task. Maybe predefined rubric algorithms would achieve what double-blind processes are supposed to and eliminate group think, logrolling, bias against negativity, ambiguity, reputation bias, and produce genuinely constructive comments in far shorter times. One sees complaints about such, and now, even more frequently it seems finding reviewers is difficult in itself especially given the explosion in the number of journals and papers being published (https://thehonores.com/3-million-papers-a-year-is-academic-publishing-out-of-control/ ). It is encouraging at least to see this possibility being explored: https://arxiv.org/html/2604.27924v1
Prompt templates can massively improve results. For instance when doing security reviews, the results are not even comparable between "Check for security issues" versus a proper template that contains checklist matrices, categories, demand to check for historical incidents, etc. Then there's another level- if you have a second agent review the first agents work recursively, it will often find stuff the first run missed.
Amen! I still file my printed emails in a filing cabinet sorted by last name. It gets complicated sometimes particularly on group email threads or when people change their last name due to marriage. But, that’s a small price to pay for having everything neat and tidy to refer back to when needed.
If the wrapper model also offers consistency and accountability, that can be worth something and could even become an enduring moat if it became a widely accepted standard. This would apply even if the wrapper was worse than a generic AI.
Such wrappers would be the missing pieces in a range of otherwise expensive and time consuming bureaucratic processes.
Arnold - beg to differ, and not because I make money differing. I've been systematically trying all sorts of tools, most of which include a variety of frontier models, but they provide the models with other capabilities, whether it is direct RAG, or SKILLs and compute, or MCPs, and so on. Refine.ink happens to be pretty helpful if you want a certain task done a certain way. So does Consensus.app, Devin, Cursor, Edison, etc. Sometimes it's a specialized sub-agent. Anyway, I also take the manuscript straight to a less specialized model, but even then, I have a pre-written system prompt (the most lightweight 'wrapper' there is, and relatively cheap).
I’ve been using a certain "wrapper" software since the 90s. It’s called TurboTax.
What it lacks as compared to the "real thing" like say Lacerte, it more than makes up for it with a tailored interface, specialized easy to understand workflows and simplicity.
In short, this post is needlessly grumpy and neglects the value of accessibility for non power users.
In what sense is TurboTax a wrapper around software like Lacerte? As I understand them, they are quite different pieces of software for similar purposes.
Lacerte is a professional grade software package that allows CPAs and tax preparers to handle complex tax situations and to go deep within the forms themselves for validation and quality control. TurboTax does some of that, but it is mostly an interview based wrapper that hides the complexities in the background and has an obvious end point on what it can handle.
Yes, that's what I understood Lacerte to be. It doesn't sound like TurboTax is a wrapper at all. Apps like Refine literally are layers on top of AI models; they don't implement the AI model, but put logic and structure around a third party model. Refine is more analogous to web hosting resellers, in the sense that the latter "wrap" the ability to host server hardware in a data center, than to TurboTax.
Based on your definition above, both TurboTax and Lacerte are wrappers in that they build logic, decision trees and calculations on top of the tax forms published by the Internal Revenue Service. In the TurboTax model, the forms themselves are hidden by default and are replaced with easy to follow questionnaires and check boxes.
That's not any definition of "wrapper" I've ever used or seen used. Implementing a standard is not wrapping the standard.
Wrapping is writing a program which asks its own questions and reformulates your data in a form more suitable for the underlying program. There is a world of difference.
Wrapping is like a CPA hiring assistants who ask a simple subset of the questions the CPA would ask, then use client answers to fill out the more complicated form the CPA uses and pass that to the CPA.
“Wrapping is like a CPA hiring assistants who ask a simple subset of the questions the CPA would ask, then use client answers to fill out the more complicated form the CPA uses and pass that to the CPA.“
I have no idea what you just said, but it sounds provocative. Is that really what you think that CPAs like myself do all day?
My guess is that pricing of top performance by big models will be such that specialized wrappers, which help minimize the cost of achieving the top performance in various legal, medical, engineering, math, programming fields will allow such specialized wrapper/ front-ends to demonstrate value added -- usually more value added than middle management.
What To Do? The key management Q. How best to do it? The engineering, workers key Q.
The middle management is to protect the decision makers from too many questions about "what", and especially to monitor the front line managers with a series of summary reports on progress towards the top decided What To Do.
Every current department of each org might be a good candidate for a very specific company - product wrapper to compare various Top ai models.
It's unlikely the best ai model will be the best in every field, all the time. For most work, the "best" answer is hardly worth more than the next best 100 answers. The best value answer is some combination of quality & price.
Quality, like in cars or anything else, has many objective features but always subjective valuations of each feature. Now I'm thinking of the classic car joke, where two guys are arguing over which car is better, based on speed, turning radius, etc. They ask a female mutual friend. She quickly answers: the Red one.
Creating a scaffold combining multiple top models to get great results seems very likely to remain valuable; and to allow a decision maker to use that ai rather than an assistant to do the comparison.
I worked for people like this: "They are the equivalent of the executives of the 1990s who had their secretaries print out their emails." A compelling analogy for folks of a certain age.
For a particular academic field, having a particular wrapper dedicated to that field gives academics collectively "someone to talk to" in terms of pointing out things that need to be fixed or, well, refined, according to the particular needs and interests that are the highest status in that field. Trying to "petition" changes in something like Claude feels like trying to petition the federal government, it's too big and too wide in scope for your complaints or suggestions to matter. But a field-specialized tool is more like a local official with whom you could actually set up a meeting. The company behind it is focusing on the opinions and usage of a particular narrow community and thus has a kind of natural closer relationship to its membership and an interest interest in paying attention to its comments and concerns and making appropriate adjustments. Claude and the others and the big companies behind them are just too big to "too inpersonal" to care about those concerns, requests, interests, and details for some tiny subgroup of its giant and diverse population of users. I can especially see the value of field-dedicated wrapper products and companies in the law. Indeed, that was part of the history of Westlaw and Lexis when the big search engines came on the scene and everybody predicted that was the end of the specialist firms. It didn't work out that way because generic search that is the 95% adequate solution for a million other users is (for various reasons I'll spare you) not a good approach to legal research, the knowledge of which is organized in a distinct kind of intellectual structure. Notice that even after several years of knowing about big problems, generic AI tools are still really bad at legal research and writing, and the legally specialist AI tools, tweaked to the particular needs of the legal profession, are (at least currently) considred clearly superior (albeit a lot more expensive, which is why lawyers try to save money and use the generic tools and are still often getting themselves into big trouble.
Using AI well is all about workflows and iteratively improved, quantitatively evaluated prompts. There is surely room for someone to come up with something much better than just feeding it straight into ChatGPT.
Whether or not Refine actually does that, I cannot say. But there is an opportunity for such a thing.
Reviewing is one of the easier, most beneficial ways that most people can benefit by using more AI.
It is great for programming, which is one of the more hyped capabilities. Even if you don't use it for generating code, having it review your patches will often give you good next steps that you would not have discovered of on your own.
It is also great for documents, though, including any form of report, email, or slide show. Similar to programming reviews, it will not just look for low-level issues such as misspellings and bad grammar, but also issues of tone and of pacing.
If you make any form of digital artifact, then having an extra reviewer tends to be valuable. If it gives you something you act on even half the time, then it's helping you produce better output, at a very low cost.
Getting back to Refine, I share the skepticism that a special-purpose tool is likely to beat direct usage of ChatGPT or Claude. The general-purpose tools learn the current best practices for basically all specialized domains. At most, you may want to add t othe prompt something like, "this is a paper intended for academic publication at (journal name)", so that it activates the relevant specialized info that it should have.
Granted, academic publication may start to really fray around the edges if the papers are generated by AI and also reviewed by them. The world is changing.
Have you tried https://econstats.org/ yet? It’s a “trusted source” tool for economic statistics
There is a popular wrapper model that uses the major LLMs for education applications called Magic School. I think much of its appeal is that it lists specific uses (draft an email to parents, design a quiz, scaffold an assignment by reading level, etc.) that AI-illiterate and uncreative teachers wouldn't think of on their own. Giving each narrow usage its own designated tool makes inputting easier for people as well.
Once AI use across education and other fields is compulsory and ubiquitous, people will probably have developed the proficiency to interact with the models themselves without the training wheels that these platforms provide.
Just to clarify, Refine has released a benchmark and outperforms 'out-of-the-box' models and open-source scaffolds by a wide margin. Creating a scaffold that combines multiple frontier models and produces excellent results is non-trivial.
https://www.refine.ink/blog/refine-ai-reviewer-benchmark
I think you should make it clear that you are an advisor to Refine. Your remarks did not seem to me to be disinterested so I was not surprised to find you on https://www.refine.ink/team. I don't have a "dog in this fight" but I think you should disclose if you do.
I thought it was obvious. But yes, I'm an advisor.
Just to clarify, I asked Claude about this paper. (https://claude.ai/share/e62ea757-35c5-4cf8-81fd-744ab703cd83). One quote: "Their scoring procedure defines a good review as one containing many verifiable, paper-grounded, atomic concerns, and explicitly "gives less credit to broad literature positioning, fit, taste, or claims requiring outside facts." But that is precisely what Refine is built to produce. So the benchmark isn't neutrally measuring "review quality" — it's measuring "closeness to Refine's theory of what a review should be." A human referee's most valuable contribution is often exactly the things this rubric discounts: judgment about whether the paper matters, whether the framing is right, how it sits in the literature. A system optimized to emit lots of granular anchored concerns will win a contest scored on lots of granular anchored concerns almost by construction.
I understand the skepticism, and there is much to improve on. That said, Refine is being used exactly for this purpose (post-acceptance but pre-publication review) by the AEA / Econometric Society. It isn't meant to replace judgement about contribution / framing.
In terms of review philosophy and the benchmark, the competing models are given a prompt that also asks for verifiable and atomic concerns. So the benchmark is not unfairly comparing review systems with differing objectives.
Writers at the post-graduate level are probably conscientious enough to use tools like this productively, recognizing the errors the AI corrects and adapting over time. Undergraduates would input their sloppy passive-voice written papers and have the AI rewrite everything without reading any feedback. This is already happening, of course.
Platforms like Refine should add a student mode, and schools should require mandatory writing seminars where students feed their papers and essays into it. An instructor would then review the results and work with students individually based on their weaknesses in writing and reasoning. It's unlikely to happen in any case because modern American universities do not engage in productive pedagogy.
Sounds like something Alpha schools would be able to do, and are likely willing to do.
Refine sounds a bit like illusion of control on the part of the Econ profession. “We have vetted this tool, understand it, and give you permission to use it”.
When in reality everyone is using AI already in their work and quite aggressively. No one is waiting around for the installed, err… elected AEA president to tell them what tools to use.
I know it’s not the focus of this post, but as someone who retired from the publishing race in economics a long time ago, I wonder: can any of these models rank the academic economics articles on the index of highest in elegance, but lowest in relevance? :)
Looking at the Refine pricing page (https://www.refine.ink/pricing ) it seems pricy, but then I have no idea of how much it would cost in terms of Claude input and output tokens (https://platform.claude.com/docs/en/about-claude/pricing ). One wonders if academia is concerned about cost efficiency?
After reading the comments, I wonder if a dedicated peer-review referee wrapper wouldn't be an ideal product even if humans are nominally uncompensated for performing that task. Maybe predefined rubric algorithms would achieve what double-blind processes are supposed to and eliminate group think, logrolling, bias against negativity, ambiguity, reputation bias, and produce genuinely constructive comments in far shorter times. One sees complaints about such, and now, even more frequently it seems finding reviewers is difficult in itself especially given the explosion in the number of journals and papers being published (https://thehonores.com/3-million-papers-a-year-is-academic-publishing-out-of-control/ ). It is encouraging at least to see this possibility being explored: https://arxiv.org/html/2604.27924v1
Prompt templates can massively improve results. For instance when doing security reviews, the results are not even comparable between "Check for security issues" versus a proper template that contains checklist matrices, categories, demand to check for historical incidents, etc. Then there's another level- if you have a second agent review the first agents work recursively, it will often find stuff the first run missed.
What was wrong about printing out your emails?
Amen! I still file my printed emails in a filing cabinet sorted by last name. It gets complicated sometimes particularly on group email threads or when people change their last name due to marriage. But, that’s a small price to pay for having everything neat and tidy to refer back to when needed.
I can beat that. I actually hand transcribe all my emails to clay tablets.
Easy when not using email at all.
Double like…
If the wrapper model also offers consistency and accountability, that can be worth something and could even become an enduring moat if it became a widely accepted standard. This would apply even if the wrapper was worse than a generic AI.
Such wrappers would be the missing pieces in a range of otherwise expensive and time consuming bureaucratic processes.
Arnold - beg to differ, and not because I make money differing. I've been systematically trying all sorts of tools, most of which include a variety of frontier models, but they provide the models with other capabilities, whether it is direct RAG, or SKILLs and compute, or MCPs, and so on. Refine.ink happens to be pretty helpful if you want a certain task done a certain way. So does Consensus.app, Devin, Cursor, Edison, etc. Sometimes it's a specialized sub-agent. Anyway, I also take the manuscript straight to a less specialized model, but even then, I have a pre-written system prompt (the most lightweight 'wrapper' there is, and relatively cheap).
I’ve been using a certain "wrapper" software since the 90s. It’s called TurboTax.
What it lacks as compared to the "real thing" like say Lacerte, it more than makes up for it with a tailored interface, specialized easy to understand workflows and simplicity.
In short, this post is needlessly grumpy and neglects the value of accessibility for non power users.
In what sense is TurboTax a wrapper around software like Lacerte? As I understand them, they are quite different pieces of software for similar purposes.
Lacerte is a professional grade software package that allows CPAs and tax preparers to handle complex tax situations and to go deep within the forms themselves for validation and quality control. TurboTax does some of that, but it is mostly an interview based wrapper that hides the complexities in the background and has an obvious end point on what it can handle.
Yes, that's what I understood Lacerte to be. It doesn't sound like TurboTax is a wrapper at all. Apps like Refine literally are layers on top of AI models; they don't implement the AI model, but put logic and structure around a third party model. Refine is more analogous to web hosting resellers, in the sense that the latter "wrap" the ability to host server hardware in a data center, than to TurboTax.
Based on your definition above, both TurboTax and Lacerte are wrappers in that they build logic, decision trees and calculations on top of the tax forms published by the Internal Revenue Service. In the TurboTax model, the forms themselves are hidden by default and are replaced with easy to follow questionnaires and check boxes.
That's not any definition of "wrapper" I've ever used or seen used. Implementing a standard is not wrapping the standard.
Wrapping is writing a program which asks its own questions and reformulates your data in a form more suitable for the underlying program. There is a world of difference.
Wrapping is like a CPA hiring assistants who ask a simple subset of the questions the CPA would ask, then use client answers to fill out the more complicated form the CPA uses and pass that to the CPA.
“Wrapping is like a CPA hiring assistants who ask a simple subset of the questions the CPA would ask, then use client answers to fill out the more complicated form the CPA uses and pass that to the CPA.“
I have no idea what you just said, but it sounds provocative. Is that really what you think that CPAs like myself do all day?