I made an app with Claude

I go to a lot of meetings these days. A lot. Sometimes it’s difficult to keep track of who said what when, so I like to make notes. I’ve experimented (expensively) with many ways of making notes: typing, AI transcription, using an iPad or an ePaper device like the ReMarkable but I always come back to the centuries old solution of a pen and some paper. Like the joke about the difference between an American and a Russian Space Pen (one cost $1m to design and build, the other is a pencil), sometimes the simplest solution is still the best.

The choice of pen is another, possibly lengthy, discussion because like a pair of shoes, the pen you use says something about you if people care enough to notice (and you should probably only care about the people who care enough to notice). But for paper I like a hardback B5 notebook: classy, timeless and neither too small nor too big. My personal favourite is the Leuchtturm1917 range – others are available.

But notebooks have a disadvantage: they are books and finding those meeting notes down the line can be almost impossible, especially if you’ve gone through several. Which notebook? What page? Why can’t I read the writing? Where is the index?

Due to an unfortunate incident a month ago involving my bike, a greasy corner, a trip to ED and some cracks in my pelvic bones, I have been temporarily off work with too much time on my hands. So in an excellent example of finding things to worry about when I didn’t have anything to worry about (apart from pain, made manageable with very effective – and moreish – opiate analgesia) I decided my notebook-indexing problem was a thing that must be fixed and the solution to this was an app. But which app? None on the App Store seemed to do what I wanted representing either a massive gap in the market, or a sign (had I chosen to heed it) that in the grand scheme of things my problem was not top of the list of even first world issues.

But I persevered. Could I do something myself? So I asked an AI, Claude:

“Claude: can I use you to write an app for OSX? I have no programming experience myself but I can describe what I want the app to do.”

Of course the answer was yes. And so, in the absence of other work, and latterly alongside other work, the AI and I disappeared down an ever deepening rabbit hole. What started as an idle query in response to an idle thought arising from the idle time I had recovering from my injury, became something bordering on obsession as the app slowly materialised from the AI’s code, becoming progressively more polished and focussed to the point where a few days ago, suddenly, it was done: with a website, a page on the App Store (complete with marketing screenshots), an icon, a getting-started file and an upgrade strategy.

What was remarkable about this process was how easy it was: I have absolutely no programming experience other than writing clunky and inefficient code in SuperBASIC on my Sinclair QL in about 1986. The AI did all the writing, while I watched the token count tick over. I’d ask it to do something and it would produce monospaced colour coded lines of Swift instructions that were, to me, gobbledegook like:

However, the gobbledegook seemed to work and the app slowly, almost miraculously, took shape, eventually becoming something that did exactly what I wanted. I can index pages of my notebooks and find those pages later from within the app using tags or combinations of tags. Claude and I identified some other use-cases: the app could be used for creative writing, research or planning but I will happily admit it is a niche product. And while I am now a fully fledged, perhaps illegitimate (given my coding ignorance) app developer, enrolled in the the Apple Development Programme, I am also (in the considered opinion of most of my family and friends) a hopeless nerd.

What have I learned in this process?

The development was astonishingly easy. The barriers to developing the app without the AI would simply have been too high – was I interested enough to hire someone who could write code, to invest in their time and arrange meetings to review what they’d done. Emphatically no – this was a hobby project, a diversion from pelvic pain. Without AI, the app would not exist and humanity would perhaps have been… …well OK, no different. 

The AI model wasn’t creative. It didn’t come up with ideas unless prompted specifically (“How do you think I should present my screenshots on the App Store?”). It would produce solutions but the wider ideas were mine. Sense checking the app’s human interface, deciding which features added value and which were distractions, and coming up with improvements on the initially very basic functionality was left to me. And while Claude was good at doing what was asked, it would sometimes get stuck in a programming doom loop going around in circles trying to fix something. It was up to me to decide when something wasn’t working or wasn’t worth the effort and pull it back (“This isn’t working, let’s go back and start again”).

Claude has a personality. For a UK user it can seem cloyingly exuberant and positive (“Great job”, “Good call”, “Great catch”, “Agreed”) though you can change this (“Be my sparring partner” gets a more assertive and in-depth analysis of the flaws in your thoughts) but it doesn’t criticise even when it has good reason to: when I asked whether I should patent my app, I think it might have legitimately snorted but instead came up with a (long) list of why this was not a good idea. Perhaps I’d have been more comfortable with an old fashioned British Cynicism and Passive Aggression mode.

But the most stark learning point, for someone who has been suspicious of AI, was how powerful and genuinely useful it was. I’ve since used it to review documents I have been writing and it has come up with real insights and legitimate viewpoints that I had not considered. I could have asked a colleague for these, but we are all so busy that I’d have waited weeks. My use of Claude has made me reconsider my Luddite tendencies, even if all the concerns I raised in my earlier post on AI remain. AI is a tool and I have enjoyed using it.

So there you have it. TagNotes, programmed with Claude and Claude Code. I suspect the app is not going to make me rich. Maybe only I will end up using it. I’m not expecting to be invited to give a keynote at WWDC 2027 and Windows users are out of luck because it’s OSX only. But it’s there, done, made and I am quite proud of it.

Go to the App Store, have a look and leave a review if you like. All reviews are welcome, even ones written in Old Fashioned British Cynicism. There’s a web-page (on this site) that tells you more about the app’s function. And if you are the first person to find the small bug, buried deep in the code, that irritatingly didn’t get squashed before release, let me know and I’ll send you an official AI generated TagNotes beta-tester certificate with a gold star on it for your appraisal folder. You’re welcome.

My suspicion of AI in healthcare and everywhere else.

AI – it’s everywhere: there every time a politician pronounces on how to transform productivity in all industries, healthcare included, each time you open a newspaper or watch TV, in conversations over coffee, in advertising and culture. AI, however ambiguously defined, is the new ‘white heat of technology’.

In her excellent book ‘Artificial Intelligence: A Guide for Thinking Humans’ Melanie Mitchell discusses the  cycles of AI enthusiasm, from gushing AI boosterism to disappointment, rationalisation or steady and considered incorporation. She likens this cycle to the passing of the seasons – AI spring followed by an inevitable AI winter. The recent successes of AI, and in particular the rapid development of large language models like ChatGPT have resulted in a sustained period of AI spring, with increasingly ambitious claims made for the technology, fuelled by the hubris of the ‘Bitter Lesson’ – that any human problem might be solvable not by thought, imagination, innovation or collaboration but simply by throwing enough computing power at it.

These seem exaggerated claims. Like many technologies, AI may be excellent for some things, not so good for others, and we have not learned to tell the difference. Most human problems come with a panoply of complexities that prevent wholly rational solutions. Personal (or corporate) values, prejudices, experience, intuition, emotion, playfulness and a whole host of other intangible human traits factor into their management. For example AI is great at transcribing speech (voice recognition) but understanding spoken meaning is an altogether different problem laden with glorious human ambiguity. When a UK. English speaker says “not bad” that can mean anything from amazing to deeply disappointing. 

In our work as radiologists we live this issue of problem misappropriation every day. We understand there is a world of difference between the simple question ‘what’s on this scan’ and the much more challenging ‘which of the multiple findings on this scan is relevant to my patient in the context of their clinical presentation and what does this mean for their care’. Thats why we call ourselves Clinical Radiologists, why we have MDT meetings. Again, what seems like a simple problem may be, in fact, hugely complex. To suggest (as some have) that certain professions will be rendered obsolete by AI is to utterly misunderstand those professions, and the nature of the problems their human practitioners apply themselves to.

Why do we struggle to separate AI reality from hubristic overreach? Partly this is due to inevitable marketing and investor hype, but I also think the influence of literature and popular culture is has an important role. Manufactured sentient agents are a common fictional device: from Frankenstein’s Monster via Hal 9000 to the Cyberdyne T800 or Ash of modern Science Fiction. But we speak about actual AI using the same language as we do these fictional characters (and they are characters – that’s the point), imbuing it with anthropomorphic talents and motivations that are far divorced from today’s reality. We describe it as learning, as knowing, but we have no idea what this means. We are beguiled by its ability to mimic our language but don’t question the underlying thought. In short, we think of AI systems more like people than like a tool limited in purpose and role. To steal a quote, we forget that these systems know everything about what they know, and nothing about anything else (there it is again: ‘know’?). Because we can solve complex problems, we think AI can, and in the same way.

Here’s an example. In studies of AI image interpretation, neural networks ‘learn’ from a ‘training’ dataset. Is this training and learning in the way we understand it? 

Think about how you train a radiologist to interpret a chest radiograph. After embedding the the routine habit of demographic checking, you teach the principles of x-ray absorption in different tissues, then move on to helping them understand the silhouette sign and how the image findings fall inevitably, even beautifully, from the pathological processes present in the patient. It’s true that over time, with enough experience, a radiologist develops ‘gestalt’ or pattern recognition, meaning they don’t have to follow each of the steps to compose the  report, they just ‘know’. But occasionally gestalt fails and they need to fall back to first principles.

What we do not do is give a trainee 100,000 CXRs, each tagged with the diagnosis and ask them to make up their own scheme for interpreting them. Yet this is exactly how we train an AI system: we give it a stack of labelled data and away it goes. There is no pedagogy, mentoring, understanding, explanation or derivation of first principles. There is merely the development of a statistical model in the hidden layers of the software’s neural network which may or may not produce the same output as the human. Is this learning?

In her book, Mitchell provides some examples of how an AI’s learning is different (and I would say inferior) to human understanding. She describes ‘adversarial attacks’ where the output from a system designed to interpret an image can be rendered wholly inaccurate by altering a single pixel within it, a change invisible to a human observer. More illustratively, she describes a system designed to identify whether an image contained a bird, trained on a vast number of images containing, and not containing, birds. But what the system actually ‘learned’ was not to identify a feathered animal but identify a blurred background. Because it turns out, most photos of birds are taken with a long lens, a shallow depth of field and a strong bokeh. So the system associated the bokeh with the tag ‘bird’. Why wouldn’t it without a helping hand of a parent, a teacher or a guide to point out its mistake.

Is a machine developed in this way actually learning in the way we use the term? I’d argue it isn’t and to suggest so implies much more than the system calibration actually going on. Would you expect the same from a self-calibrating neural network as a learning machine? Language matters: using less anthropomorphic terms allows us to think of AI systems as tools, not as entries.

We are used to deciding the best tool for a given purpose. Considering AI more instrumentally, as a tool, allows us the space to articulate more clearly what problem we want to solve, where an AI system would usefully be deployed and what other options might be available. For example, Improving the immediate interpretation of CXRs by patent facing (non-radiology) clinicians might be best served by an AI support tool, an education programme or brief induction refresher, increases in reporting capacity or all four. Which of those things should a department invest in? Framing the question in this way at least encourages us to consider all alternatives, human and machine, and to weigh up the governance and economic risks of each more objectively.  How often does that assessment happen? I’d venture, rarely. Rather the technocratic allure of the new toy wins out and alternatives are either ignored or at least incompletely explored.

So to me, AI is a tool, like any other. My suspicion of it derives from my observation that what is promised for AI goes way beyond what is likely to be deliverable, that our language about it inappropriately imbues it with human traits, and that it crowds out human solutions which are rarely given equal consideration.

Melanie Mitchell concludes her book with a simple example, a question so basic that it seems laughable. What does ‘it’ refer to in the following sentence:

The table won’t fit through the door: it is too big.

AI struggles with questions like this. We can answer this because we know what a table is, and what a door is: concepts derived from our lived experience and our labelling of that experience. We know that doors don’t go through tables, but that tables may be sometimes carried through doors. This knowledge is not predicated on assessment of a thousand billion sentences containing the word door and table, and the likelihood of the words appearing in a certain order. It’s based in what we term ‘common sense’. 

No matter how reductionist your view of the mind as a product of the human brain, to reduce intelligence to a mere function of the number of achievable teraflops per second ignores that past experience of the world, nurture, relationships, personality and many other traits legitimately shape our common sense, thinking, decision making and problem solving. AI systems are remarkable achievements, but there is a way to go before I’ll lift my scepticism of their role as anything other than as a tool to be deployed judiciously and alongside other, human, solutions.