8 ms·
Show HN: Penny-1.7B Irish Penny Journal style transfer
Yesterday, in the bygone hour of the weekend, I undertook a most singular and fascinating endeavor, wherein I delved deep into the recesses of my mind, and, with a fervent zeal, breathed life into a most remarkable creation. I embarked upon the quest, with the singular object of fashioning an artificial construct, one imbued with the verdant essence of the Irish Penny Journal, an ancient and venerable tome that holds within its pages the whispered tales of a bygone era.
In my haste, I set forth to construct a dataset, a repository of those fleeting moments, these ephemeral sentences, which spoke of a bygone age. I procured a collection of these fleeting moments, these sentences, and with them, I synthetically conjured forth modern translations, an ingenious feat of substitution, which allowed my artificial construct to take on the guise of the language of the Irish Penny Journal.
Then, with great anticipation, I fashioned a small encoder, a humble instrument, with which to guide the artificial construct in its endeavors. I presented this encoder as a bribe, a reward, to a most ingenious system, one that trained a colossal language model, one of unbridled potential, one that was capable of weaving tales with the very essence of the Irish Penny Journal.
And lo! In the succeeding moments of time, I witnessed a most wondrous thing. My artificial construct, armed with this training, and guided by the whispers of the encoder, began to speak, to speak in the language of the Irish Penny Journal. The words it spoke were, indeed, the words of the past, imbued with the nostalgia of a forgotten era.
And thus, my friends, I have witnessed a most singular creation, one which embodies the language of the past, yet, in its most recent iteration, speaks to the present. A testament to the ingenuity of the human spirit, this artificial construct speaks of the bygone era, yet, with each word, it whispers to us, to us, of a future yet to come.
——
That’s Penny explaining itself to you. This was trained using GRPO only, in less than a day using a single A6000. I didn’t use any SFT, and only relied on a small encoder (MiniLM2) trained to classify texts from the Irish Penny Journal and their modern translations (synthetically produced).
- ekianjo 1y agoNice work ! It still manage to use the word 'delve' in the first sentence, which is a giveaway that it's written by a LLM.
- deepsquirrelnet 1y agoHaha! You’re right. Maybe I’ll add a penalty for that and some other giveaway words and make a revision.
- dragonwriter 1y agoIts considered an LLM tell because its a term that is rarely used by the median modern casual writer; its not at all uncommon even in current literature [0], and its even more common in older literature, so complaining about it in a model designed to reproduce a particular style of 19th century print literature is silly in the extreme. [0] many “LLM tells” fit this pattern of just being common features of professionally-published works that are less often seen in casual writing.
- observationist 1y agoIt's very common in fantasy novels - dwarves and wizards do a lot of delving into caves, dungeons, and towers. It's also a solid academic term, so scientists delve into a lot of subjects, and brain people do a lot of delving into psyches. LLMs are raising the bar by expanding the vocabulary people are exposed to, so words like delve will stick out. I think it's preferred by writers because it articulates a nice sounding alternative to words like explore, venture, analyze, think about, confront, etc. It's a useful, versatile word, and one of the metrics by which writers measure quality is the minimization of syllables. LLMs are mostly indistinguishable from humans at this point; a one-shot output from any of the major models can be recognized in the same way you might recognize a writer. With multiple style passes, you're not going to be able to tell the difference between ChatGPT, Ronald Reagan, Bill Clinton, Hunter S. Thompson, Einstein, or any other sufficiently modeled figure. Throw in a few tens of thousands of words written by yourself and most of the models will do a nearly flawless job of copying your stylometric profile.
- bee_rider 1y agoDelving also has the implication of going deep, while exploring has the implication of going wide. I wonder if human authors really do a good job of picking the right word there. Either way I wonder if “misused delve” could be an interesting signal.
- veggieroll 1y agoHave you written anywhere in detail on how you gathered your dataset and trained the finetune? I have a few use cases that are like this, but I'm not sure where to start.
- deepsquirrelnet 1y agoMy dataset is here: https://huggingface.co/datasets/dleemiller/irish_penny_journal https://huggingface.co/datasets/dleemiller/irish_penny_journ... It’s fairly simple — I essentially just split the original text into chunks and then used some bigger models on openrouter to clean it up and provide translations to modern English (seemed to be pretty easy for an LLM). After that, I just trained a MiniLM2 model to classify the texts. I used this in a reward function for reinforcement learning and changed the system message as a simple instruction to write in the prose of the IPJ. I debated whether or not to use any SFT, and decided not to. I think if the style would be too hard to learn you might need some seed/cold start SFT data. I’ll try to get my scripts up in github for you to look at. It’s just a few short training scripts.
- veggieroll 1y agoThanks for the explanation! I'm learning and and think this would be a good next project for me to try, especially since I have a real world use case in mind with a similar amount of data available. In particular, I'm not very familiar with reinforcement learning and am not sure how you use the embeddings from MiniLM2 as a reward function. (Edit: maybe this is the jaccard similarity?) I'd really appreciate it if you were open to posting scripts! I see a few snippets around and could probably cobble something together after a while. But, it's cool to see something already working to make sure I'm not getting too far off into left field.
- deepsquirrelnet 1y agoYou can ignore the jaccard similarity field. That was just to monitor the text->cleaned text conversion to make sure it didn’t stray too far from the original while it was fixing whitespace OCR issues. I didn’t use embeddings. Nreimers account on huggingface has the minilm models which are BERT-like, but trained using distillation. https://huggingface.co/nreimers/MiniLMv2-L6-H384-distilled-from-BERT-Large https://huggingface.co/nreimers/MiniLMv2-L6-H384-distilled-f... Is the one I started from. You can then just load that and train it on your data using a standard transformers classification pipeline. ChatGPT can zero shot that part reasonably well if you gave it this description. From there you should check out the GRPO trainer in TRL. It has taken me a bit of time to learn how to use it effectively. There’s a TON of parameters in the configuration, and occasionally I have to hunt down arxiv papers to understand them.
- bee_rider 1y agoIt is sort of funny that the Irish ended up being the best practitioners of the English language, despite the fact that they were forced to use it.
- blululu 1y agoNot sure this is true. Most of the famous Irish writers were Anglo-Irish Protestants (Yeats, Wilde, Swift, Beckett). Joyce is the notable exception here. The Irish certainly produce great cultural works of the English language (well beyond their size). But also the penal laws greatly depressed the cultural output of the Irish people for 250 years.
- projektfu 1y agoI feel that the "best practitioners" is not limited to the most famous writers. A great thing about Ireland is the conversations to be had there, and how quick-witted Irish people often are, with clever use of the language. This can be true elsewhere in the English-speaking world, but Ireland has some renown for it.
- deleted 1y ago[deleted]
- w10-1 1y agoJoseph Conrad also had English as a second language.
- _1 1y agoKinda of strange to pick an example that is just wrong. It's supposed to be written from 1840 and says Paris is the seat of Napoleon almost 20 years after he died.
- Philpax 1y agoIt's transferring the style, not the knowledge.
- deleted 1y ago[deleted]
- sjkoelle 1y agoMarvelous! What gain beyond zero-shot would motivate a humble citizen to implement this instrument? How was the superiority assessed?
- deepsquirrelnet 1y agoGood question - my best assessment is just the text classifier. IE was the LLM able to “trick” the classifier into believing the text came from the IPJ? And it came quite a long way in training. Initially the classifier scores were very low (mean around 0.05, meaning modern). Over training, the scores came up and ended close to 0.95 (IPJ). The standard deviation of the group also declined, so the consistency of responses improved as well. My thought on the application of this is that you could use it to create different voices to your responses and probably even add multiple at a time to a single model. I chose this one to experiment, because it is easy to classify and the data was available in the public domain. GRPO kind of opens up RL to lower tiers of hardware and I’ve been able to experiment with it at home. I think this is something people can do themselves and it’s fun and potentially useful in games or possibly in relation to applications interfacing kids with lower reading levels (eg using a reading level classifier instead).
- dwringer 1y agoYet, one might justly question the imperative of cultivating a distinct model for such an endeavour, when a judiciously framed prompt, enriched by apposite examples, might suffice to imbue a sophisticated engine with the desired stylistic graces. Though it is undeniable these modern engines shall wax greatly in their proportions, and the art of discovering the exact prompt to elicit their most felicitous expressions is a task far from trivial, yet, it must be admitted, the pursuit holds a certain diversion for the inquisitive mind! It is, perchance, not the creation of manifold engines, but rather the artful disposition of singular contexts, that shall bestow upon diverse interlocutors their proper and unique voices.
- kamranjon 1y agoThis is really cool! Do you have any of the pipeline code available that you used for training? I am curious about how you created the reward model. I love little projects like this, thanks for sharing. I've been fine-tuning on my mac and an interested in getting into GRPO, which I haven't tried yet.
- deepsquirrelnet 1y agoI put my scripts up in github. It’s a bit scrapped together at the moment — but fine as a reference. https://github.com/dleemiller/PennyLM https://github.com/dleemiller/PennyLM
- latchkey 1y agoReminds me of this: https://www.unix.com/man_page/debian/6/jive/ https://www.unix.com/man_page/debian/6/jive/
- joshstrange 1y agoNow I'm just imagining a video game with characters each having their own fine tune applied on top for their dialog. I'm guessing you could use some relatively small models. In each case you would be feeding all the context to the model (player name, current relevant quests, summary of previous interactions, etc). Though maybe fine tuning/training isn't even needed and a good enough prompt will work (Not sure what all they used for this [0]). I'm excited for the first AAA game that tries this. Anyone that has played a RPG-style game knows that after a few times going into a city (or a couple play-throughs) the dialog feels repetitive. I love the idea of Skyrim but with better dialog. You could either run the models on the user's computer or maybe just run it on the backend so you can block certain generations (wrong/misleading/"unsafe") and just ship updated dialog lists to the client occasionally. [0] https://www.youtube.com/watch?v=d6sVWEu9HWU https://www.youtube.com/watch?v=d6sVWEu9HWU
- jsheard 1y agoCounterpoint: NPCs repeating their dialogue serves as an implicit indicator that you've exhausted their content and it's time to move on. If they gain the ability to make vapid smalltalk forever then you'll forever be second guessing whether you're wasting your time on them. (also spare a thought for the poor QA testers who would be given the Sisyphean task of making sure an LLM dialogue system always stays in character and doesn't hallucinate non-existent or outdated content/lore/mechanics)
- joshstrange 1y ago_Very_ good point. I had not fully considered that, same deal with conversation trees vs free-form entry/response.
- acdha 1y agoIt’s a really good point. One thing which comes to mind is the way some games distinguish between UI blocking dialog and background color, which could be a great place to start: imagine walking through a city like Baldur’s Gate only it has actual thousands of people who are saying different things when you walk by, and some of those are based on things your party has done recently with specific details about appearance, gear, and actions which would be too hard to do with traditional dialog approaches (e.g. kids talking about a battle and who they thought was best like real kids talk about sports, a good priest wondering what a paladin was doing spotted talking to a notorious thief, etc.). Something like that could add color and immersion without affecting gameplay or wasting anyone’s time, and you could extend it to things like vendors (“saw you put that axe to good use…” or “were you wearing these boots when you freed those slaves? I bet my brother will want buy them!”) to flesh out the approach before using it for load-bearing purposes.
- fitsumbelay 1y agothis is awesome
- KaiserPro 1y agoI'm not sure if you've tried this already, but removing the translate step might give you a more authentic output. In the journals that I saw, the language was much more simple than the output.
- throwaway314155 1y agoYou mention no supervised finetuning. May I ask why? I'm curious if you could get similar/better/worse results by just finetuning the LLM on your dataset rather than generating synthetic data, training a classifier and using GRPO? Cool stuff in any case.
- deepsquirrelnet 1y agoIn some other experiments, I’ve noticed that SFT can be very rigid and hard to generalize from, and later found this paper: https://arxiv.org/abs/2501.17161 https://arxiv.org/abs/2501.17161 Also I wanted to start from the instruct model, and wasn’t certain if it would be a good idea to do continued pretraining over top of it. Otherwise I’d need to create an instruct dataset for which the passages from the IPJ would be the answer… that also seemed a bit unnatural. Perhaps there’s still some viable pathways in there, but I decided to see how the stupid simple thing worked out first and just go directly to RL. I think supervised fine tuning is feasible, but it’s not entirely straightforward what to train on and how much data to use if you want to follow up with RL.
- npunt 1y agoLove it. Immediately reminded of the text filters back in the day like the pirate one that would drop letters and replace with apostrophes and change certain passages into "arr" or "yarr matey"
- sterlind 1y agowhat a wonderful work of whimsy! well wrought. I'd love to have a library of these, so I could pipe text into `penny`, `brainrot`, `pony`, `newspeak`, `corporate`, `scp`, `trek` etc. have you published the training notebook somewhere?