The AI tools I dropped this year, sorted by why
# The AI tools I dropped this year, sorted by why
A screenshot reading tool I paid for spent three minutes on my fleet dashboard and handed back a tidy table of numbers, correctly formatted, none of which were on the screen. The demo I watched read a restaurant receipt flawlessly.
Receipts are a solved shape. Someone has thought hard about receipts. Nobody has thought about forty rows of small grey text with a status dot on the end.
I have a piece about the AI tools I actually use. This is the other half and it carries more, because the keepers are short and boring while the graveyard is large and the reasons in it repeat. I have also written about the automations I regret building, which covers things I made. These are things I adopted, and they fail differently, because somebody else decided what the tool was for before it reached me.
Nothing that lost gets named
Sorted by reason instead of by tool. The reasons generalise and the tools do not, and something that died on my inputs may be your best purchase this year.
For context on whose inputs are judging: mobile proxy lines on carrier SIMs, a rack of Android phones rented by the week, a video pipeline running across fifteen brands. So the real files here are Singaporean accents, telco names, dashboards nobody designed for a machine to read, and logs.
Reason one: the demo was a chosen input
One transcription tool gave me a clean, well punctuated transcript of nine minutes about SIM provisioning in which a Singapore telco appeared as five different words, none of them the telco. Its demo was two people with studio microphones talking about podcasting.
A demo is a chosen input, chosen because the tool is good at it. The choosing is invisible, and the choosing is the story.
So the filter costs ten minutes and happens before I finish signing up. Ugliest real file I own goes in first, never the sample. About half die in there, and I would otherwise have paid two months for each, because I am slow to admit something is not working.
Reason two: it worked, and it wanted me to work differently
A note taking app, a properly good one. It wanted every idea and clip and link to land inside it, and in return it would join things up for me later.
I used it well for eleven days. I was enthusiastic in a way I now read as a warning.
Then a customer line went down on a Wednesday and everything I wrote that day went into a text file on my desktop, because that is what my hands do. I never opened it again.
The tool did nothing wrong. What beat it was a habit that costs nothing, is already running, and is about nine years old. The same thing killed a project tracker and an AI email client.
This is the position I would defend hardest. Most tool adoption fails on fit with an existing habit, and capability barely enters into it. A vendor can do nothing about that from their end. They can make the product better. Familiar is not a thing they can ship.
Which is why two honest people review the same tool in opposite directions and both are right. The review describes a fit, and the fit lives in the person.
Reason three: absorbed by something I already paid for
The largest category by count, and I did not judge these at all. A summariser I paid for over seven months stopped mattering when the model I already pay for could hold a whole document in one go. A diagram tool went the same way, and a code search product lasted until my editor shipped the feature.
Each was a good decision on the day I made it and stayed good for eight or nine months.
What changed is what I sign. If a product does one thing, assume the one thing becomes a checkbox inside a bigger product within a year. Single purpose tools get paid monthly now, whatever the annual discount is. That discount bets the feature survives, and I have lost it enough times to stop taking it.
Reason four: the price moved while I was three layers deep
One tool removed the tier I was on and the next one up cost several times the money for capacity I would never touch. Another moved from a flat monthly fee to metered, which at my volume is a different product wearing the old name.
I have no idea what their costs looked like. Inference is expensive and a lot of early pricing was a guess funded by somebody else's money.
The money was the small part. The expensive part was that the thing sat three layers inside a pipeline, so leaving meant a rewrite.
Everything I pay for per call now sits behind one function in one file. When a model I depended on was retired the swap took an afternoon, because of that file. A vendor buried deep in a pipeline you wrote two years ago has stopped being a supplier and become a landlord.
Reason five: good tool, wrong person
I bought an evaluation platform once. Well built, and built for a team where several people have to agree a model change is safe.
I am one person. When I want to know whether a change is better I run twenty rows whose answers I already know and read the misses. An afternoon of work, in a file.
I did not buy it because I had the problem. I bought it because I wanted to be the size of business that has the problem, which gets said out loud almost never.
The two I want back
Here is where I found out my judgement is tilted.
Two years ago I tried a tool that kept versioned prompts alongside the outputs they produced. I dropped it inside a week. It added a step to a loop I ran forty times a day, and the payoff was a graph I could not read yet, holding four days of history.
Eighteen months later I spent an afternoon building a worse version of the same thing, because I could no longer answer why a prompt had changed and when. Mine has no interface and I maintain it.
The second one: I wrote off small local models after a single bad week. They were losing at a job that I now know was the wrong shape, because I had one prompt doing four things at once and no model of any size does that well. I came back a year later, split it into three calls, and models that size have handled two of the three perfectly ever since. I blamed the model for a design decision I had made.
What joins those two took me long enough to see that I am embarrassed about it. I judge a tool on the friction of the first hour, and the first hour is loud and visible. On anything that accumulates, the value shows up around month six and it is quiet. So my judgement leans toward keeping whatever feels easy today and cutting whatever pays off slowly, and every tool I regret dropping sits in the second group.
The same tilt shows in the keepers. All of them were useful on day one. That says something about me and little about whether they are the best available.
Two clocks, not one
Tools that replace a task get an hour, today, on the worst real input available. Transcription, generation, conversion, lookup. If it fails there it will not improve, because the failure is about your data, and your data will not change to suit it.
Tools that accumulate get a quarter. Notes, evaluations, logs, prompt history, test sets. These are empty on day one and empty is the design, so day one tells you nothing. Judging one in a fortnight is judging a garden in March, and a fortnight is roughly a free trial, which is not an accident.
The cost is that I keep things for three months that I would have killed on day four. The two above are what I paid to learn that.
Paying rent on the corpse
The opposite failure is holding a tool you should have cut, and the tell is in how you describe it. If your answer is the setup, the templates, the weekend you spent learning it, you have recited an inventory of your own spending with the tool absent from it.
One question does the work. Name the last decision this changed, with a date, inside the last month. If the honest answer is a plan for using it properly once things calm down, it is dead already.
I am bad at this. A plan on my card renews until March 2027 for a service whose paid path I stopped using in about April, with a quarter of a million credits unspent. I use the same vendor daily on the free path, which at my volume does the identical job. Cancelling takes ninety seconds. What stops me, I think, is that cancelling makes it official, and while it renews I get to tell myself I might go back.
Where this is weak
I almost never revisit. I do not read changelogs and I do not return to a tool that failed me once, so some of the above is wrong right now and I would not know. The transcription one especially.
I also drop things silently. I have never written an honest reason in a cancellation box, so no complaint of mine reached anybody who could act on it. Which makes me part of why the category stays as it is.
And the fit argument turns back on me. My habits are nine years old and nothing about them is sacred. Some of what I call a bad fit is me being unwilling, and unwilling looks exactly like unable from where I am standing.
Get new guides and videos first — join the Telegram channel.