Claude Sonnet 3.5 was my personal inflection point for LLMs. It was the point where I recognised the potential and started using the models. One of my early uses was as an editor for this blog.
LLMs are, and continue to be, outstanding at picking up poor sentence constructions and bad grammar. But something about the other advice, about how my writing was structured and flowed, never struck me right. My opinion is that my writing is better without it.
LLMs keep evolving and so I experimented with using them for writing an unpublished piece.
The prompt
My system prompt is fairly simple:
You are a blog editor.
Your role is to provide feedback around flow, prose & content. Particularly identify areas where the meaning might be difficult to decipher.
This blog is in British English.
You cannot provide examples of specific phrasing.
At the end, in a simple list, will you point out any grammar & spelling errors.
If there are any image links in the text, ignore them. It does not matter that you cannot see them.
To talk through some decisions there:
- I specifically try to focus the prompts on areas of my writing that I consider my weaknesses.
- Removing specific phrasing from the output ensures that quirks of LLM writing stay away from my writing. Of course, this is imperfect. We are all consuming LLM writing at record rates and we will gravitate towards using words we read.
What is better
LLMs no longer need to be guided towards offering a breadth of feedback. For instance, if I am writing an argument, they will critique the argument, despite the fact that the prompt did not instruct them to do so.
LLMs can generate an infinite amount of critique for an argument. As someone who has the tendency to take a good idea, and apply it too widely or without the correct caveats, this can be a useful tool to point out flaws I hadn’t noticed.
The models no longer compliment me relentlessly. They are much narrower in their appraisal of my words.
What is still bad
The same model, with the same configuration, with the same input, will give you two different opinions on the piece. One might find your argument quite persuasive, while another will be completely unconvinced. One might convince you to insert a piece of prose, while another will convince you to axe it.
Judgement, as ever, is still important. It is clear that the value of a brilliant writer will win over amateurs like me, even with an LLM to my rescue. I’m not convinced there is much value for those with talent.
The opposite side to the infinite critique is the inability to satisfy the machine. The feedback treats you as an opponent in a debate. If you’re writing in Pinker’s classic style, this doesn’t really work for you. I’m not aiming to present at a debate club.
Models have a tendency to find a golden argument in your writing, and proclaim that this should be your locus. Except, if you do follow the advice, the model will pick up another point and say that this should be promoted to the top. This is a never-ending cycle.
The expert problem
Fundamentally, this comes down to an inherent problem with LLMs. They will generate all manner of output, but to use it effectively requires the touch of an expert.
The expert is needed to guide the model towards the correct output1. Moreover, they’re needed to know what output to discard and disregard.
I am an amateur writer. My highest qualification comes from secondary school. What gives me the qualifications to judge an LLM’s editorial feedback? While I may think my prose is above amateur level, that is my own judgement.
Which makes me feel this entire endeavour is problematic. Arguably, some pushback is better than none, but what if the bad pushback drives my writing to be lower quality, and less digestible? What if I encode these bad principles permanently into my writing, and thus see a permanent quality drop?
I think I’m going to stop using it, again
LLM writing is universally a struggle to read. It has no panache to make it worth reading. Ergo, refining your own writing too much in the style of an LLM will do the same to your own.
It is possible that this is a steering problem. A better system prompt might create feedback that is more representative of how I want to write. But again, what is the right feedback?
For now, I am going back to trusting my own flawed instincts. At least they’re mine.
Footnotes
-
Terence Tao has a different ChatGPT to me. ↩