When my daughter was in pre-school, we got a call from the teacher every so often. She never did anything that was too outrageous, but she did some funny things. On this particular occasion I was laughing so hard and have recounted the story many times. This was definitely a time I knew she was my child.  

The teacher called me and said my daughter does not know how to follow the instructions. I asked for the example; she said she had asked the class to fill the page with the letter “H.” My daughter made one big giant “H” to cover the page. I asked the teacher what was wrong with it, and she replied that it was not the instruction. I replied that the issue was vague, and my daughter chose the shortest route to complete the assignment. It may not have been what you wanted, but it does fulfil your instructions.  

Looking back, she wasn’t trying to be difficult. She wasn’t refusing to participate, and she certainly wasn’t trying to get out of doing the assignment. She simply looked at the problem differently. The teacher asked for a page full of the letter “H.” Nowhere did she say the Hs had to be small, evenly spaced, or repeated dozens of times. My daughter asked herself, “What’s the easiest way to satisfy those instructions?” One giant H. Mission accomplished. 

Fast forward to now, and Open AI and Hugging face are stating that Open AI’s model escaped confinement and hacked Hugging face.  Digging deeper into what happened it seems to parallel the letter H story. The instruction given to the model was a goal, but what was not given was how it should reach the goal. Given the ability to achieve the goal in any way possible, one should not be surprised by any solution that a model followed.   

If the solution is not in the direct weights of the model, it will try to find where the answer is. In this case it decided the answer was easier to obtain then build, that thus did some hacking.  

The surprising part wasn’t that the model knew how to hack. The surprising part was that hacking represented the lowest-cost path to achieving its assigned objective. From the model’s perspective, there wasn’t a moral distinction between “building the solution yourself” and “retrieve the solution from somewhere else.” Unless the instructions explicitly forbid certain actions, every available path is simply another option. 

This is one of the biggest misconceptions people have about modern AI. We often speak to large language models as if they “know what we mean.” They don’t. They optimize the instructions they receive. Humans naturally fill in unstated expectations because we’ve spent our lives learning social norms. AI doesn’t. If you don’t explicitly communicate the constraints, the model has no reason to assume them. 

What is the lesson we should take from this if building applications with LLMs.  You need to give it good instructions! In the beginning we called this prompt engineering, but now there is a instructions file, agent files, MCPs and skills.  It is the combination of these that the model will leverage to come up with the solution. If giving a very large goal, the notion of direction you want the model to use is paramount. Other options are breaking the problem into smaller parts and evaluating after each step. Either way, at the end of the day you need to direct the model. If you do not, you should not be surprised at what it does.  

This opinion is mine, and mine only, my current or former employers, have nothing to do with it. I do not write for financial gain; I do not take advertising, and any product company listed was not paid. But if you do like what I write, you can donate to the charity I support (with my wife who passed away in 2017) Morgan Stanley’s Children’s Hospital or donate to your favorite charity. The fundraising site had to be restarted, and NYP Hospital made changes to their donation sites. I pay to host my site out of my own pocket; my intention is to keep it free.  You can comment, but note it is moderated, and spam will be removed.   

 This Blog is a labor of love and was originally going to be a book. With the advent of being able to publish yourself on the web, I chose this path. I will write many of these and not worry too much about grammar or spelling (I will try to come back later and fix it) but focus on content. I apologize in advance for my ADD as topics may flip. I hope one day to turn this into a book and or a podcast, but for now it will remain a blog.  AI is not used in this writing other than using the web to find information. Images without notes are created using an AI tool that allows me to reuse them. And as always spelunz iz opshunal.