← Back to the CV

Essay

Explorations with AI

Having worked in the software space for nearly two decades, I was curious to find out what AI does to this area and also how it might shape my future in this industry. Instead of borrowing the opinions of others I decided to run a months-long “experiment” to form my own. There were two sets of objectives, first was obviously to gain hands on knowledge of working with AI agents. Second was to answer a few specific questions:

The outputs of the experiment are currently deployed in app stores. Links below should you be interested to see them. All free.

For the overarching experiment I decided to borrow a way of working from motorsports, the regulations generally describe a test that is run to verify specific parts. As long as the part under scrutiny passes the test, the part is deemed legal to use even if other teams perceive it to have an unfair advantage.

So one important restriction that I put on myself for the purposes of the experiment was to view the end product i.e. the usable software and not the raw code at any point.

To run the experiment, I decided that as ostensibly AI is supposed to augment my capabilities I would concentrate only on the big picture and not go into the minutiae. Hence I limited myself exclusively to conversing with the agent and not working in an IDE. I made sure that the agent organised my thoughts and wrote extensive documents recording each feature that I wanted, along with why. Every subsequent decision that I had on the features or the overall direction also had to be recorded along with the reasoning behind it.

I blindly accepted the agent’s implementation while judging the working end product. Essentially treating the agent like a capable developer which initially had no idea of why I wanted the functional requirements done in certain ways. Regarding architectural decisions I asked the agent for its opinion on why it designed certain things in certain ways just to gauge if its reasoning made sense, gently pointing it to the best practices that I was aware of instead of letting the agent overcomplicate it as it is wont to do.

In short, I put on my PO/BA hat and let the agent be the dev and architect. Here are the results of the experiment.

Can I build something end to end using AI and have it be usable?

Short answer, yes. Here’s proof:

Pellucid (Mac, Windows)
Mac App Store · Microsoft Store
Kalkra (Mac, iPhone, iPad)
App Store
GastRotator (Android)
Google Play

These are built for me, or my near ones. These are usable for their intended audience, your mileage may vary.

What additional tooling and infrastructure are required to make sure that that process is repeatable and consistent?

This is the tricky part. Short answer is: a lot.

My breakthrough came when I mentally categorised AI agents as brilliant but with memory going back just a few minutes. Most of the custom tooling that I had to build was to let the AI store why I asked it to do something. Storing the reasoning behind a decision enabled it to come up with better quality output, even after a few days and numerous new sessions. Also shared memory to use across different projects regarding the coding standards or principles that I wanted the agent to adhere to, UI/UX guidelines and lessons learnt in solving specific problems.

Keeping daily logs helped immensely as well, git history is good, but having a short note in plain text explaining when something was done sped things up a lot.

Finally and most importantly, the biggest breakthrough was wiring up a comprehensive test suite that tested every feature that was added. Full TDD, tests written first and features written to match them.

Surprisingly enough, using skills and guardrails did not benefit me enough. The light bulb moment was when I realised that anything important enough for the AI to do always could not just live in an agent.md file or a skill.md file, it had to become infrastructure rather than instructions.

Deterministic tests for everything, be it complex mathematical solvers or simple checkboxes or validating approved colour combinations. The biggest problem by far that I found in the experiment was that the agents kept changing code that had nothing to do with what I had asked it change. And I was running around clicking around finding out what else the agent changed (remember the parameter that I had set saying that I won’t review the code before the merge).

The decision to use a very pedantic TDD approach saved the experiment, otherwise it may have failed to deliver any usable products. TDD wired in as a pre-commit CI gate, with minimal output indicating which feature failed, enabled the agents to self-correct before presenting me with the outputs.

One interesting note here, some agents were better at following TDD while others reported that they were following TDD but routinely “forgot” to do so. Hence the pre-commit gate. Some agents “cheated” TDD by designing tests that were simply placeholders. Only one agent family reliably did the tests hence it’s that one that I feel the most comfortable working with. Others needed separate sessions to fix the tests themselves. These separate sessions were needed, again, as I refused to look at the code and the tests myself.

Which part of the development lifecycle does AI impact the most?

Given a fixed token budget, is it better spent on code generation, or automated testing of the code, or making the detailed plan for each feature to be built?

Short answer: unprecedented freedom to experiment with different solutions.

The one thing that was holding me back was the traditional process of writing individual “stories” to hand over to the agent. I simply could not match the speed at which these were delivered and initially I was finding it difficult to reach my weekly token quota in the first place. This is an interesting finding from the experiment that I feel highlights most the capabilities of AI right now. Traditionally teams were limited in output by the total amount of developer bandwidth available. Now the bottleneck shifts left to the requirements phase itself.

This however opens up the possibility for experimenting with different possible solutions before committing to one in the final product. This goes hand in hand with the ability to cheaply and deterministically test for each and every tiny thing to ensure that the solutions are truly viable with proof.

My experiments point that if you have enough time and money, having AI work beside you in all three areas from requirements to implementation to quality control is viable. But if I was on a limited budget I would spend a large portion of it letting the POs/BAs experiment with prototyping proposed solutions.

To put it simply, as agents can code a feature quickly and without a lot of time required by a dedicated developer, the cost of failure is drastically lower and that changes the consequences of being wrong and opens up the team to more experimentation before committing to a solution.

Some teams spend a lot of time in meetings exploring what a feature should be, where it should be, would it be useful at all before a developer is first asked to have a look at the feature and weigh in. What that means is that by the time a developer actually starts work on a feature, days and in some cases months have passed from the time someone mentioned the feature in the first place. Ergo even a small feature is expensive in terms of overall time taken to get from concept to prototype. This makes the decision of killing a not so useful feature more difficult and time consuming than it should be.

With AI however, getting from concept to functional prototype now is almost a quick solo job. A PO or a BA can get that done without having to interact with anyone else. The prototype can then be demoed and accepted/rejected without much baggage attached to it and purely on the merits of the feature itself. The only caveat being they have to have a fully functional sandbox to play in.

Having a functional prototype also gives the dev a head start in implementing the feature as they can see the expected outcome and don’t have to interpret what the final result would do based on documents/stories handed over to them.

Where should I invest my time in learning to stay relevant in the “AI-first” era?

To throw in some jargon, I believe this is the right time to develop the so called “comb-shaped skills”. AI means that a PO or BA can also build and test. So having better understanding of exactly what makes sense to build (e.g. should this be a function in the front end itself or should we put it in the back end where it can be reused across services) or what to test or indeed if the DB is normalised enough for the task at hand, I feel are skills that the humans working should develop in addition to just generic AI literacy. In other words, if you know some details of areas adjacent to your area of expertise, AI can contribute some of the details and that results in the entire team moving forward as a whole.

Before attempting this experiment, I had gone through a bunch of courses on how to use AI, and there was very little from them that helped me at all as they covered what was state of the art in 2024 or 2025 which as it turned out was not as relevant in 2026 as the frontier labs had already solved the issues from the prior years. So I feel the only way to stay relevant would be to use it daily and in areas that you work in and immediately adjacent to them so that you can judge the output yourself and take action to correct it when the agent makes a mistake.

Conclusions

What I keep hearing from industry leaders when they speak about AI is that they are very cautious about the governance associated with using AI, with the implication being AI is only for developers for now, with the leaders being afraid of having another shadow IT/EUC problem. Based on my experiments, I feel the answer is giving people on the product side sanitised sandboxes to play with AI and build simplified business processes, which then, if found to be useful, can be handed over to the architects and devs who have more knowledge of what makes the system resilient and fits the current best practices in the org.

With the bottleneck shifting from development to requirements, giving the business domain experts the ability to build using AI speeds up everyone as they focus on the functional requirements while the developers can then focus on the non-functional requirements and architectural problems. Having access to a very capable AI agent also means that it’s not really a big divide between FRs and NFRs as either group of experts can increasingly contribute more to the other, thereby increasing the overall velocity at which products are brought to their intended users while retaining the capability to deterministically ensure the quality of the codebase in a way that was unthinkable earlier.

A different way to put it would be, I feel, in the future the agenda of the meetings would shift from:

“Here is a story that needs implementing, let me know if you need a walkthrough.”

to:

“Here is a functioning version of what I mean, let’s figure out if and where it belongs in the production systems.”