This text was generated using AI and might contain mistakes. Found a mistake? Edit at GitHub
Building Software in the Age of AI — a Conversation with Randy Shoup
A note from our team.
Software-Architektur im Stream will be streaming live from the ISAQB Software Architecture Gathering in Berlin this November.
Join us on site.
You’ll find more about the program and special discount code for our community on our website software-architektur.tv.
Hello and welcome to another episode of Software-Architektur im Stream.
Today with me is Randy Shoup.
Randy, may I ask you to introduce yourself?
I hope I pronounced your name the right way.
Thank you.
Actually, it’s a German name.
And so as we were chatting personally, it’s Shoup.
Even though it’s spelled incorrectly, we can have that conversation later.
Ralf, it’s great to be with you.
So yeah, hi, I’m Randy Shoup.
Right now I’m the head of engineering for CircleCI.
So we’re a continuous integration service provider, as many of you will know, through my career.
So I’ve been doing software professionally for about 38 years.
So quite a long time.
During most of that time, actually more than half, I called myself an architect in one form or another.
So I’ve been chief architect at eBay.
I’ve been chief architect of a security company.
And whether my title had architect in it or not, I always have loved software architecture, distributed systems, and so on.
So again, I worked at eBay, I worked at Google, and I worked at a bunch of other Silicon Valley companies.
So I hope we’ll talk about all those things.
So what you didn’t say, but what I found out through Claude, that yeah, even as a small child, you already had contact to the newest technology at Xerox Palo Alto Research Center, Xerox PARC.
So you were raised with technology, the newest technology.
So yeah, it makes sense that you worked at those huge companies.
So how was your time when you played around?
I know we didn’t talk about this, but now that I…
Yeah, yeah.
To have the newest technology, did you know that this was really something which others didn’t have access to?
I did not.
So the brief story is, my father earned his PhD in computer science in 1970.
And he earned it at a university in the United States called the Carnegie Mellon University.
It’s a pretty good computer science university.
And his was one of the first PhD programs in computer science in the United States.
So obviously people had been doing computer science for decades, but this was the first time it was a serious field of study.
And he and my family came out to the San Francisco Bay Area, which became Silicon Valley later, to start to do that work.
And my father joined the Xerox Research Lab there in 1971, I think.
So not immediately after his PhD, but one year afterward.
And he was in the very first group of researchers that joined there.
And so other people that joined around the same time include Alan Kay, who invented Smalltalk, of course.
Butler Lampson, Chuck Geschke, who founded Adobe.
Let’s see.
The guy who founded Microsoft Word, whose name is escaping me at the moment.
Simonyi, Charles Simonyi.
Anyway, a whole bunch of great people.
And the serious computer science in the early 1970s was very small.
Everybody knew one another.
I can tell you a story about meeting Don Knuth, the very famous professor who my dad and his friends knew really closely.
They were joking at each other at this later time, which I could tell you about if that’s interesting.
Anyway, but you asked the question, what was it like as a child to have that experience?
Yeah.
I mean, when you’re a child, you have no idea that your experience is unique or common.
And so I didn’t really fully understand until much later in life that it was very it was weird for me and my brother to beg my dad to take us to work on the weekends, which we absolutely did, because when we were very little, we would play in the beanbag conference room.
So I had a conference room without no chairs and tables, but just beanbags.
And we would make forts and play as kids do.
But then my dad’s research was in computer graphics, and he built one of the very first paint programs.
So if you think of Mac paint or Microsoft paint right now, all those metaphors were stuff that he and his colleagues invented in the early 1970s.
So having one part of it, that’s the canvas where you draw, and another part, which is the palette where you choose your color and your brush, you choose your color and your brush, and you go and you draw on the canvas.
And so only later do I realize that my dad’s research at Xerox PARC was the only research that would be interesting to a six-year-old, because they did other interesting things that, you know, the first word processor, they did the laser printer, they did Ethernet, they did object-oriented programming, they did touch screens, they did VLSI logic, you know, which is a way of laying out circuits on a chip.
I probably missed one, but like, they invented that group at Xerox PARC in the early 1970s, invented multiple decades of computer science right at that time.
But it was lovely.
I really enjoyed drawing spaceships on my dad’s paint program.
So you grew up in the middle of the heart of technology where everything was invented, and I guess there were moments where you noticed that some technology now reached the consumer state where you said, hey, that was something I played around with some years ago.
Yeah, it was really.
So there’s a much longer history, obviously, associated with Xerox PARC, but just very briefly, the goal of the lab was to build the office of the future and essentially build a personal computer.
Now remember, we’re talking 1970 to 1973 here.
The first, the IBM PC didn’t come out till nine years later, 1982.
The Mac didn’t come out till 1984.
And it was really only the Mac in 1984 that had the same combination of things that they had developed at Xerox PARC.
Again, graphical user interface and WYSIWYG, there you go, WYSIWYG work, essentially, you know, what you see is what you get.
And that was the main line of research there.
And again, the laser printer was part of that.
Ethernet was part of that.
Object-oriented programming was part of that.
And the graphical user interface is part of that.
My dad’s research was unrelated in computer graphics.
But the Alto, which is the machine that they build at Xerox PARC, basically became the Mac.
And there’s a wonderful story about Steve Jobs getting a tour of the lab in 1979.
And by some accounts, stealing the work.
By other accounts, being inspired by the work.
It doesn’t really matter what, how you want to talk about it.
But the, but it is a direct line in our industry from the research and the office of the future and personal computing at Xerox PARC into the Mac and then Microsoft Windows and then how we use computers today.
Fascinating.
So you experienced all of this and then did your own way work for large IT companies.
And now we have the age of AI, if I might say now.
So for, I think, three or four years, we now experience all the LLMs, how good they are.
So it started out with LLMs where many people said, Hey, it’s giving me the wrong answers.
I can do better development than I get those, yeah, those auto corrections and so on from the AI.
And it evolved.
And so now how do you experience how the AI LLMs evolve?
What’s your current take on it?
Is it really already helpful?
Is it really AI or is it still the, yeah, how would you say, the stochastic parrot?
Yeah, stochastic parrot.
That’s such a wonderful phrase.
And I, obviously it’s not my phrase.
Is it AI in the sense of, is it human level intelligence?
No, not yet.
I have lots of thoughts about AI.
And I like to think that I’m pragmatic about it rather than being completely pro or completely anti.
I think several things.
Number one, AI as a tool is absolutely transforming our industry.
There’s no argument about that one way or another.
So we can, I think you cannot be a professional software developer and completely ignore AI.
You can use it more or less, but you can’t ignore it.
It exists just like electricity exists, just like compilers exist, just like a bunch of other tools that we have in our world.
It is very clear that the mechanism of typing in code is, AI is able to do that.
There’s no argument in my view that AI can take, and again, through experience, AI can take vague ideas or even better, very clear specifications and turn them into code.
That used to be a large part, maybe the whole part of what it meant to be a software engineer.
And so I definitely understand for me and for everybody how our jobs are changing.
That is not all of software engineering.
It was never all about the typing.
And we are experiencing the industrialization of software engineering, which we had never had before.
So just like, you know, in the industrial revolution or the agricultural revolution, the evolution of technology has made it so that we used to do things manually by humans, and then we have machines to do part of the work, not all of the work.
I think we’re experiencing that as software engineering, right?
So the LLMs are able to do a lot of the work that we used to do manually, typing in code.
Does that mean that humans and architecture are not useful in software engineering?
It doesn’t mean that at all.
At the same time as AI is doing all the typing work for us, at the same time, what we need to engineer is the system to make that safe.
What do I mean by that?
AI LLMs produce code at a rate no human can keep up with.
What is now the engineering problem is not the typing.
AI can do the typing.
The engineering problem is for us to make that safe, and that is every aspect of the evaluation of AI-produced code.
So that is static analysis, that is linting, that is using static typing so that you can check things, the compiler can check things, all the way up to having adversarial LLMs review, you know, LMA produces code, LLMB reviews it.
That is spec-driven development, where we say as a spec what we want the LLM to generate, and we check the spec via tests and via evals at the other end.
There’s a lot more to that eval, but I’ll just stop there.
I think AI is a good producer of things, and now the engineering problem is channeling and almost back pressure against the LLM generation.
But you just, yeah, named all those things we learned during the last 20-30 years how to do good software engineering, and somehow it seems it was lost.
Not everything was used in software development, and now we notice that, hey, it might be a good idea to write a good proper spec and to have some tests and static code analysis to work with the LLM together.
Isn’t it funny?
It’s funny and also wonderful, because LLMs are telling us in a new way all the things we knew about building good software.
It’s forcing us to build software in the correct way.
And to make it clear that I’m not missing anything, that is, to your point, test-driven development, behavior-driven development, where we specify the behavior in an executable form, that is, writing a spec ahead of time and validating that at the other end, that is, small iterative forward movements instead of let’s type for six months and then think we’re done, produce incremental value and check that incremental value or that iterative step along the way.
Literally everything that we know, we, this architecture community, have known maybe frustratingly for decades about how to build good software, AI is reteaching us that.
Why?
Because we have, how could we get away with doing software without these verification steps?
It’s because we were implicitly doing it in the software engineer’s brain, right?
The software, I mean, whether the software engineer was actually writing the tests or not or doing it in a formal way, the only way we got away with it was because we had a human in there and the human could manually, in some way, manually do the verifications that we’re talking about, right?
And obviously the more we let the computers do it, the more we use tests, the more we use static code analysis, the more we used compiler stuff and metaprogramming, the easier and easier.
But AI is making, AI gives us the power of producing the typing without the human judgment.
And so we need to encode that judgment explicitly into the evaluation part.
Does that make sense?
So yeah, it makes sense.
Someone in the chat writes, he or she has always distinguished between the role of a developer and that of a programmer.
I think when we had punch cards, it was even more easier to distinguish.
So the AI is now the programmer and we are the developer or something like this.
So we try to get a big picture and think about it and give the task a purpose.
The task which the AI fulfills, we give it a purpose.
Is it something we can say this way?
That aligns with how I think.
I agree completely.
I don’t think it’s, whether we call it programmer, developer or whatever, I will say this.
So I agree with the sentiment and the words don’t really matter.
The AI, so it is now free or near free to produce code.
Does that mean we’re done?
No, because does the code even work?
Does the code do what we think it does?
Does the code break everything else?
Does the code violate invariance?
Does the code violate the architecture?
So all these questions are now explicit things that we need to encode in an automated way.
Why in an automated way?
Because LLMs produce code at machine speed and we need to be able to evaluate that code at the same machine speed, number one.
Number two, I know everybody knows this.
How did LLMs even learn?
It’s feedback, right?
It’s RLHF.
It’s like a, forgetting my words, but like, because I’m tired, it’s a learning loop.
And the way that LLMs learned anything is through this loop.
And the way that they’re going to learn how to do our stuff is a loop that we need to produce, right?
So we need to be, so as LLMs produce iterative new codes, new PRs or whatever, we need to have evaluations that are right there, giving that LLM direct feedback, both to stop the LLM from doing things wrong and violating stuff, but also to teach the LLM to do it better the next time.
But the LLM is already trained.
I can’t really modify the weight.
So the teaching is in terms of setting up a coding loop.
So I noticed that, yeah, it doesn’t stick to my linting rules.
So I will give it a linter so that it can check its rules.
Yeah.
That will always be true.
So as the models, the LLMs are getting better and better, of course.
And also, no model knows exactly our code.
And by the way, I don’t want to go too far along this, but like, let me take a step back.
Every time we do software, we’re doing something new, by definition.
Yes.
Because unlike agriculture or industrial engineering, we’re not building the same thing over and over again.
It’s free to copy the bits.
So by definition, every time we ask a human or an agent or an LLM or whatever to build a new piece of software, it’s a new piece of software.
So the LLM is trained on more software.
It’s read all the books, all of them.
It’s trained on more than any human could ever be trained on.
And also, that’s still not enough, right?
We need to, as I was trying to say, we need to encode all of the concerns that we care about.
So again, static code analysis, compiler checking, tests for functional adversarial evaluation, security vulnerabilities, architecture.
I’m missing some, but like everything that we care about that we would do in a human code review, we want to encode that.
Oh, and here’s what you were saying.
The A way to do that is to continue to try to teach the LLM upfront not to make mistakes.
We should do that.
And also, that’s never going to be enough.
So we always have to have an independent evaluation of what it does.
And we can do that independent evaluation with other LLMs, right?
So an adversarial LLM A produces some code, LLM B checks it.
We should, we can and should do that.
But also an even more efficient way to do it is to, every time we can make a deterministic, not a deterministic check, a static code analysis, a traditional code way of evaluating something that is both more deterministic, so more likely to be correct, and also way cheaper, like by orders of magnitude cheaper.
It’s like, you know, LLMs cost like a human in some sense, the cost, right?
So like there are a bunch of things that we don’t ask humans to do.
Like we produce chips in fabrication facilities.
Every one of those, every single chip that’s produced by Intel or Samsung or TSMC has to be checked, but not by a human.
Yes, but when you talk about those checks, I mean, some parts of those checks obviously belong to the harness.
They can be executed in a very fast way.
The LLM already has a compiler to have syntax checks, and some other parts belong to the build pipeline, because they take longer and are not executed with every request.
So it seems that you are currently at the heart of this development, evolution, how we handle those things.
And I find it fascinating to see what part should be in the harness, and how we speed things up, so that we recognize when to run what kind of test to speed those pipelines up.
Yes, fantastic question.
Let me see if I can, I’m going to restate it slightly differently, because it’s part of the answer.
Our goal is ultimately to produce customer value, and the way that we do that is producing good software.
How do we produce good software?
As you were talking, I can think of at least four points.
I want to type it right the first time.
Yes.
Right, so that is context engineering.
That is a good LLM model.
That is context engineering.
That is a good spec.
Probably other things.
Type it right the first time.
Then there is, does it compile?
So there are things that, like turning the typing immediately into a binary, there are already things that we’re checking in there, which we always have been.
Then, right after that, there is what at CircleCI and other people in the industry are calling the inner loop.
I’m going to talk about the inner loop and the outer loop, which is kind of what you said.
So, after, did we type it, did we type what we wanted the first time?
Did we check it via compiling in that exact moment, maybe linting as well?
And then there’s the other part of the inner loop, which is what are local tests I can run and local evals that I can run very, very fast at machine speed, right?
So, at the same speed as I can produce code, I can do those evaluations.
And to your, that’s the inner loop, all of those together.
What you point out correctly, and this has always been true, and it’s still true, is there are things I want to check that are too expensive to check in that inner loop.
That’s the outer loop.
And the outer loop, we call that CI, and like, don’t do it for me, but like, even if I didn’t work at CircleCI, there always was and is, today in software, that outer loop where I’ve done as much, as efficiently and economically as I can, I’ve checked my work as much as I can in this inner loop, and then now there’s more tests that take longer or take more resources or, you know, for various other reasons, need to run in a special environment that need to run in that outer loop.
And you ask a very good question, which is, where do the evaluations run?
And the answer, I can give you the hand-wavy answer, which is, as many as possible should run in the inner loop, as can be efficient and economic.
And then, if it’s not economic there, and it also matters, you do it in the outer loop.
And the reason why I give you that hand-wavy answer is because what is and isn’t economic or efficient is actively changing.
Do you see what I’m saying?
Yes.
Yes.
I feel what you’re saying.
For me, the outer loop is a slow one, but every developer has the same build loop.
So build pipeline.
And that’s a benefit.
And with the inner loop, the local build pipeline, I get a feeling like snowflakes.
Different developers do it in a different way.
So on a team, you have different quality of source code, which enters the build pipeline.
It might even be something that one developer tells the system, hey, remember not to make this mistake anymore.
And it lands in the agent’s MD of this one developer.
And so, yeah, you get those snowflakes, which are slightly different.
Yeah.
Yeah.
I agree.
So we’re going to violently agree about this.
To your point, as good as we can make the inner loop, we still will always need the outer loop.
Right?
I agree with that.
We need the outer loop for several reasons.
To your point, my inner loop might be different from yours, and so maybe we don’t trust that they’re the same.
That’s legitimate.
Also, to your earlier point, there are checks that we have to run that are not efficient enough, fast enough, efficient enough, economic enough to run in that inner loop.
But the other thing is the build pipeline and that outer loop need to also have a slightly different goal.
And that goal is producing code we’re actually going to ship or execute.
And it is correct for that outer loop to have, because it has a broader requirement, if that makes any sense, like it has more things to do.
But what we’ll see over time, and these are all compatible ideas, what we are seeing and we will continue to see, in my experience, is we will be moving left as many evaluations as possible.
And I’m probably pointing in the wrong direction from you guys.
But like the outer loop, think of it as a, you know, it’s a line of a workflow.
And the earlier and earlier and earlier we can do an evaluation, the better.
That’s for sure true.
Right?
Because we get better feedback.
And it’s easier to modify when it’s earlier in this pipeline.
So that’s why my answer, you know, you were like, when do we do which checks?
Like do it as much as you possibly can in the inner loop to the left.
And there will always be more that you’ll want to do, particularly because you’re building an artifact at this point, you know, in the outer loop.
And like that requires another whole set of checks on that exact artifact.
So what I do understand is that the outer loop is also some kind of alignment layer.
So your code will be checked the same way as my code.
And so the linter will test the same rules for both of us.
And if, for instance, I don’t use a linter and just let the code generate whatever is the LLM generate whatever code it likes, then my LLM will hit some boundaries in the outer loop, which is it has to correct, will run more correction runs than your LLM, which already has a linter, or maybe it was already trained in a different way so that it doesn’t need the linter anymore.
It just writes the code as intended.
Yeah, and let me, I think we’re violently agreeing and I can see in the chat that people are agreeing as well, which is great.
Let me just add some numbers to this, which I happen to now know at CircleCI.
So CircleCI does, is again, we’re a CI provider.
I’ve been there for all of two or three months, but CircleCI has been around for 14 years or so.
And probably there are customers here on the call.
We do every year, a research study, which is called the state of software delivery.
So the state of software delivery report, you can go Google it and download it right now.
I’ll tell you about it though, because it’s super interesting for this conversation.
Because CircleCI runs a lot of people’s build pipelines, we get to see a lot of build pipelines.
And what we see is that there is a huge difference between the companies that are moving the very fastest from even the average company.
And the difference is not the fastest companies are moving twice as fast.
They are moving nine times faster than the average.
I’m not talking about, it’s not nine times versus the worst.
It’s nine times versus the average, versus the medium.
That’s fast.
It’s really fast.
What in the world, and your next question I hope is what in the world are those great companies doing that makes them nine times faster?
And per people in the chat, it is engineering the build pipeline so that you reduce the opportunity and the impact of merge conflicts.
So what we see in the average is, I’m forgetting the numbers of this, but it’s in the report.
I’m making this number up.
Maybe it takes two or three or four tries for a change to pass CI and get into the production or get into the artifact.
I forget the number.
Please, we’ll look it up later.
And what we see is that number is much lower, meaning there are much fewer iterations at the CI level for those companies that are moving nine times faster.
Does it make sense what I’m saying?
So they are moving faster exactly because they have structured their work and their pipeline so that it flows.
There’s not a lot of back, I mean, like I’m thinking water pipes, right?
There’s not turbulent flow, which is exactly the engineering term for it.
There’s not like I add my change and it comes back and it comes back and it comes back.
They are able to get changes into production more effectively with fewer reiterations at the CI level.
And again, there’s a lot of ways to do that.
And we know that they are doing all of these things, the companies that are really good.
Number one is the more that you can modularize your code back to things we already knew about software, the better this is all going to be, right?
Because I’m working in my little area and you’re working in your little area and they don’t conflict with each other because we’ve architected it or designed it in an appropriate way.
The other thing they do, like we were talking about before, is moving as many checks left as possible.
It’s not like these best companies make fewer, produce fewer bugs or have fewer mistakes.
I’m sure they produce just as many mistakes or maybe more than the average.
But they’re detecting and resolving those mistakes earlier in the process before it gets to the CI part of it.
Could it be if they managed to detect it earlier and also streamline the process so that not as many defects are surfaced, that they also don’t need so many manual interventions and they trust the process more?
So I guess if you don’t trust it, you have to manually review everything and get slower.
And if it always hits some error correction loop, you wonder, hey, what’s going on there?
Let’s have a look at it.
Yeah, I don’t know this for those particular companies, but this isn’t part of the study I’m referencing.
But also in the next couple of weeks, a industry white paper about AI code review is coming out.
And I and eight other people are co-authors on this.
And again, what we find, independent of my experience at CircleCI, but what we find is the companies that are leveraging AI the best and the companies that are moving the fastest are exactly because they have engineered the harness or engineered the eval loop super effectively.
And it’s not completely taking humans out of the loop, but it is instead everything that you can validate via a computer, you validate via a computer, and only the things you can’t validate with a computer do the humans do.
You see what I’m saying?
I’m saying back what you said and in just slightly different words.
And again, as I was saying, yeah, I was going to say two things.
Number one, evaluate as much as you possibly can with a computer instead of a human.
So every time we have something that we care about, does this follow our architecture standards?
Does this follow our coding standards?
The more that we can encode that in software, either LLM or even better, something deterministic, every time we discover a thing that slipped past, we engineer it back into the pipeline as opposed to just checking it by a human over and over and over again.
And a really inspiring talk that I heard a couple of weeks ago, unfortunately it was a private, was about software is not new in this idea.
And so two industries are really very similar to this semiconductor manufacturer and pharmaceutical manufacturer.
I’ll tell you why I say that.
Again, we produce chips and we’ve been producing semiconductor chips for decades.
There is a trillion dollar industry.
I’m not lying.
Trillion dollar industry called semiconductor test equipment.
Wow.
This is not producing the semiconductors.
This is producing equipment that tests semiconductors.
So companies like KLA and Applied Materials, and I’m sure many others that I don’t know off the top of my head, each of which are like, that’s 200 billion, that’s 200 billion.
Like you add them all up, it’s a trillion, more or less a trillion dollars of enterprise value in machinery that validates that semiconductors we produce through these fabs are correct.
This is very similar to this is the same engineering problem I hope we’re seeing as software right now.
We are right now building the semiconductor test equipment industry, if you see what I mean, to help with A.I.
Right.
A.I. is the fab and and our harnesses and our evals are the semiconductor test equipment, similar in pharmaceuticals.
So, again, the goal of a bunch of pharmaceutical manufacturer is so once you discover a drug, you have to produce it reliably.
And produce it reliably means these are biological things like they can get infections and they can be broken in chemical or biological ways.
And like it’s it’s human level dangerous.
It’s mission critical to make sure that when we’re trying when we produce a drug or some substance at industry speed that we don’t kill people.
And so there’s another trillion dollar industry that’s about I forget the phrasing of it, but essentially pharmaceutical test equipment.
So we produce this batch of I don’t know, what’s the new fancy thing, the diet drug or whatever, like we produce a batch of this like, you know, substance and we need to validate at machine rate, at industrial rate, whether that is the is the correct substance.
And there’s again, there’s a whole trillion dollar industry that is an engineering industry around validating that the pharmaceutical outputs are.
That’s quite interesting, because when you tell about talk about this, it sounds that these industries have developed one set of tooling, one which which I can just buy to test something.
And when I think about software development, I just asked Claude, what do I need to to set up my testing?
And it comes with I think it was 50 tools, something like this to to have everything covered.
And wouldn’t it be a great idea to have something just here you go.
That’s it.
That’s the whole harness you need as a whole pipeline and just run it and it will work.
I wonder whether we will, yeah, develop something like this in the industry.
Yeah, I think we’re in I think, you know, so I thought you were going to go a different direction.
I agree with this also.
I thought you were going to go in the direction of those disciplines are actual engineering and we are now learning how to do engineering in software.
Yeah, yeah, yeah, yeah, absolutely.
Yeah, I mean, there are disputes about whether software engineering is engineering.
So it’s it’s more it’s more and more engineering every year.
Like it’s not inherently not engineering, but it has been this is not even wrong.
It has been a craft for many years.
I am going to answer your question in a moment, but let’s riff on this for a moment.
For for most of human history, making textiles, clothing that you wear is something that you did manually within your own family.
Similarly, furniture like for most of human history, if you had something like whatever you lived in and whatever you slept on was something that you manufactured locally and manually.
And then very recently in human history, you know, maybe let’s call it 400 years or so for textiles now, 150 years, let’s call it for now maybe 200 years for industrial things we’ve made.
We’ve industrialized both of those.
Right.
We’ve industrialized textile manufacturer and we’ve industrialized machine and furniture manufacturer.
That takes nothing away from the craftsmanship that went into initially making furniture or making making clothing.
But but but in order to do those things at scale, they had to transition.
And we’re in exactly that state in software right now.
So software is going has been for maybe a number of years and more and more with AI is going from it’s a craft again, not wrong.
That’s beautiful craftsmanship to something that’s industrialized.
Right.
Go ahead.
That’s quite interesting, because I think about now software, the term individual software is somehow changing because now everybody uses LLMs to to create some web pages to.
So it’s back to the family where it’s crafted before industrialization.
I don’t know.
I really like this.
Oh, wow.
Yeah.
I really like that model.
Yeah, I mean, now I’m now I’m this is going to be a great conference talk.
I’m not going to give it.
But from craftsmanship to industrial to mass produced industrialization to industrial scale personalization.
Yes.
So that’s what what it looks to me like now.
But there’s still the risk of white coatings that, yeah, you have the wrong harness.
So if we industrialize the harness, then everybody can create its own furniture as someone likes.
Yeah.
Yeah.
That was the question that part of your question.
That was the interesting part that I didn’t answer.
I think I think that is coming.
And I think we will in the same way as.
Compilers and compiler tool chains were used to be very bespoke, very unique, very almost individual, like different teams like you and I are old enough to remember if you ever did any Windows development or Microsoft development, like the first thing you do is you write your memory manager.
And then and then you write like I hope there are some nodding people listening as well.
So you write your memory manager.
Then you write all your like data structure implementations, right?
Like nobody in there.
No, but I mean, nobody, five people in the industry are like some very, very small number of people are iterating on memory managers.
And those are very, very smart people.
And I respect them.
And there are very few people.
Few people, but thank goodness they exist that are iterating on data structure implementations.
And we’ve moved past that in our industry by standardizing.
I mean, depending on whether you’re in dot net or like there’s a standard set.
I know you know this.
I’m saying things we already know.
Yeah, I mean, of how do I do memory management in my environment and how do I do data structure and basic algorithm stuff in my environment?
Those come for free.
But what you’re asking for is exactly that with this eval stuff, right?
Like we’re in the we’re in the exploratory phase where what’s the way to say this?
It is one right now.
One could go out and assemble a world class set of evaluations.
But I would have to accept like we could we would have to assemble that.
And what you’re saying is I’d love to get that assembly out of the box.
And I agree.
I remember the time where everybody had to write his or her own string class because it was better.
It was more performant.
Yeah.
So nobody does this anymore.
Yeah, yeah.
And yeah, so again, just to say just to put the things together, I think that I think we are right now in the write our own string class phase of developing these evals or these harnesses.
Yes.
Everybody writes their own harness and says, yes, my harness works.
That’s on the shoulder and doesn’t notice that the LM got better and now creates already what we need.
And the harness, the work the harness has to do is getting less and less.
But we are reaching out to other evals, which we notice that, yeah, we need and can be done.
Yeah.
Yeah.
Just to just to say that I’m so strongly agree.
And let me say that in a more in more general terms.
Let’s see, there’s a wonderful framework from Kent Beck, which he calls the three X framework.
It’s explore, expand, extract, explore, meaning try a bunch of new things, expand, find one or two of those things and really scale them.
And then extract is like taking profits in some sense, like using that as the baseline for the next for the next thing.
So explore, expand, extract.
And we’re in that explore phase, meaning we’re as an industry, we’re trying a lot of different things in to evaluate LM output.
And the next phase, which we’ll have all together, is the is the standardizing around some small number of them and then, you know, leveraging leveraging them.
Similarly, in machine learning, there’s explore-exploit, right?
So like, try a bunch of machine learning.
I’ve been doing machine learning for 25 years, whatever, doesn’t make me an LLM expert, that’s for sure.
But I do know a little bit about machine learning, as I know many of your listeners will, and explore-exploit, rather.
Yeah.
So we now talked about a lot of things where we do agree, and what we talked about with the Harness and the Ebers, it sounds for me like a decision whether, yeah, do it by ourself, or wait for a product to hit the market, and just flow with the market.
What do you think?
What’s better for a large company at the moment?
Just use the tools which are on the market, and yeah, wait for them to get better.
I mean, each month they are getting better.
So yeah, that’s a good question.
I don’t think I have a good general answer, honestly, because this is very context-dependent.
I will say, what’s the way to say this?
Everybody who produces software with LLMs needs evals, that’s for sure, even without LLMs.
Yes.
But particularly, again, back to our story about LLMs are forcing us to do good software.
We already knew how to do good software, and now we are absolutely forced to do good software if we’re going to leverage things at machine speed with LLMs.
I think right now, I could, and you could, and any company could assemble a set of evals with a bunch of off-the-shelf things, whether I buy them from a vendor or use open source.
I think we could everybody do that right now.
I actually do think, I think if we’re going to use, I’m convincing myself, if I’m willing to use LLMs to produce code in my company, I also have to be willing to spend a lot of time on that evaluation.
That’s just a fact.
And I think you can’t have both.
You can’t say, I want to produce code at LLM speed, and I’m going to wait to evaluate it.
Right?
Yeah, yeah, yeah, yeah.
But that’s why I get a feeling many companies still have a foot on the brake, because they say every line of code has to be reviewed manually, and because they don’t have the evals and don’t, even if they have, they don’t trust them.
Yeah.
Well, I’m sure that, well, that is true.
That’s absolutely true.
And what I, I’m going to answer that in a second.
The overall, I don’t want to give people the wrong impression.
The overall industry has a very wide distribution of AI usage, very wide.
So there are those companies that I mentioned that are 9x, do have 9x, 9x faster, fewer merge conflicts, more flow than even the average.
They are on the, you know, big users of this.
Yes.
Still, I’m making these numbers up, but that’s maybe 5%, 10% of companies, something like that.
And there’s a long, long, long tail of us who, not even wrong, like have not fully adopted AI, like we’re not industrialized.
And like, we shouldn’t feel bad.
That’s just a fact of the industry.
And by the way, in a company that’s very AI forward, like where I work right now, the individual developers, we have this massive, an individual team.
So like we have several teams, which we call software factory, and they’re very small and they use AI stuff and evals, and they’re generating mostly Greenfield, lots of new, good work.
And then still in our company at CircleCI, the vast majority of engineers are leveraging AI in their day-to-day work for sure, but are still working on building the eval harnesses that we’re talking about.
So again, I want everybody to just like not feel bad if you’re in the large part of the distribution rather than this small part, but that, you know, this part’s growing.
Okay.
So you’re asking what should a company, what should a company do?
And is it, what’s the way to say that?
You’re asking, I’ll frame your question differently.
How can we use leverage LLMs to build at machine speed and evaluate that building at human speed?
You can’t.
I’m not saying give up, but like you can’t do that.
So what should you?
So, well, what we used to do is we built at human speed and we evaluated at human speed.
That worked, kind of, right?
But you cannot build at machine speed and evaluate at human speed.
So what you have to do is all the things we’re talking about is take those humans that instead of evaluating, instead of having every human evaluate every line, take a few of those humans and engineer with the help of things that are all the way out in the industry, engineer a pipe, a pipeline, an eval set, whatever word you want to use, a harness to evaluate that.
So what I’m not saying is, and then walk away and it’s machines all the way down.
What I am saying is engineer the harness exactly like semiconductor test equipment manufacturer, exactly like pharmaceutical equipment manufacturer.
Do that and then you will have machine production of code and some amount of machine evaluation of code.
Now what’s left, that is okay for humans to evaluate.
Does that make sense?
Yeah, it makes sense.
So in my own words, I would state something like it makes sense to have a team of toolsmiths who build the harness in order to enable the rest of the developers to go faster because they can rely on this harness on the evaluations.
I think not every developer should create those evaluations by him or herself.
So this team can then, when they notice that tools are out there, use them.
You need the machine speed for the reviews and somehow you have to create it and decide how.
Sorry to interrupt you.
I get excited.
We’re finally agreeing.
And one of the users typed in the engineering challenge moves from code to harness.
Beautiful.
Wow.
Couldn’t have said it better myself.
Absolutely beautiful.
And that’s what we’re trying to say is you can’t, again, it’s just a say back, you can’t do machine speed code production and human speed code evaluation.
Those don’t work.
And that doesn’t mean humans aren’t useful.
That means humans step one level higher in the abstraction and engineer the evaluations and the specs up front and, you know, but engineer the whole process.
And to your point, which I think you made very well, should that be every single engineer of one thousand or ten thousand?
Like, no, that’s not efficient for other reasons.
Not because they couldn’t do it, but like that’s duplicative of effort and that causes its own problems.
So take every company at large scale these days that I’m aware of that’s doing this well has some kind of platform team developer experience.
You can use lots of different words for that.
I ran this team at eBay when I was chief architect.
And that is the team that engineers the production of code for the company.
By the way, I know everybody knows this, but I’m going to say it out loud.
We already need, before AI, we already needed that team.
And now we really, really, really, really need that team.
Or we need the help of that team.
Right.
Again, just like LLMs are forcing us to do good software development techniques that we’ve already known.
Similarly, LLMs are forcing us to do good, I don’t know, software development organizational techniques like, OK, if everybody’s doing the same thing or roughly the same thing, let’s extract that out into a common service or a common team or a common platform and do that.
That’s just efficiency.
But, you know, the problem I do see is that engineering is fun.
Now, if engineering is now more in the harness than in the coding, I mean, the harness doesn’t produce value in the term of, here’s a product which I can sell.
It produces value in terms of, yes, here is more quality, like the test system in the semiconductor industry, where, yeah, you only have a few companies who produce it and everybody uses it.
And the value is in the different types of semiconductors.
And so I wonder, I think we will not need so many harness engineers.
So the fun jobs in harness engineering will not be so many.
Oh, I don’t know.
I have several thoughts.
We there’s a there’s a thing called the Jevons paradox.
Yes.
And you’re familiar.
But just to say briefly for those who are listening that aren’t, Jevons was an English economist in the mid-1800s and he was talking about coal.
But what he noticed, which was very surprising, and this is the paradox, is that when coal extraction became more efficient, it was cheaper to get coal.
Somehow in aggregate, Britain spent more on coal.
Wait, what?
You made it cheaper to get coal.
And yet in aggregate, we spent more.
Why?
Because coal now when coal was whatever, 10x cheaper.
Now it became things that used to not be economic to do with coal now are.
Same with software.
So let me finish my thought, please.
Just two seconds.
And I know you know where I’m going.
With LLMs, we have reduced the cost of one part of software, which is producing the software.
That has gotten cheaper.
So one part is the engineering problems move upstream and downstream.
Upstream, what are we building?
What matters to customers?
And how do we write a spec that validates that?
And then it moves downstream to engineering that the typing that the LLM did is correct.
That’s one force, that the engineering moves away from the code upstream to what are we doing for customers and downstream, how do we validate what we did is correct.
Additionally, Jevin’s paradox says software became cheaper.
We’re going to do more software.
That is absolutely the fact.
And to your earlier point, maybe a lot of that software will be for individuals or maybe it’ll be for companies or whatever, but guaranteed.
So Jevin’s paradox requires two things.
One is a resource gets substantially cheaper.
The other part is that the demand for the resource is what’s called elastic.
In other words, there’s no limit to the demand for this thing.
And so like energy, there’s no, sadly for the planet, there’s no limit to our demand for energy.
There is a limit to our demand for food, for example, like there’s only so much food I can eat.
But cognition and software, there’s no limit to the demand for software.
There’s always more things to automate, more things to make better, more things to make more efficient.
Anyway, so as much as LLMs are changing the day-to-day life over time, are industrializing and changing the day-to-day life of a software engineer, there are many very important engineering challenges and also times many more things we’re going to do with software.
So I’m not scared for our profession.
I think there will be more engineers tomorrow and the next year and the next year than there ever were.
Sounds good.
And yes, I think this comes also again down to, am I a programmer?
Do I just type the lines of code or am I a developer who solves problems?
And we will have problems also in the future.
So that’s not…
Yeah, that distinction is wonderful.
I’ve heard that framed a bunch of different ways.
I love the programmer developer one.
That’s really nice.
Another way is, are you motivated by the journey or are you motivated by the result?
Right.
So, and there’s no, one is not better than the other.
Like it’s okay to be, it’s okay to have like the typing.
I really did like the typing.
It’s really, it’s fun.
It’s problem-solving.
And then the other related one is, um, I think something is wrong with your microphone.
It’s, oh, I’m sorry.
Can you hear me now?
Can you hear me now?
Yeah.
Yeah.
Now it’s, uh, there, there, there are some disturbances.
So yeah, let’s, let’s, let’s we’re, we’re about to end.
But anyway, the, um, the other distinction that you raise is craftsperson versus being an industrial engineer.
Yeah.
Um, so we now agreed on lots of things and we are already shortly over time.
I would like to go back to the start where you said, yes, it’s AI, but it’s not like human intelligence.
And we are in an industry where we invented duck typing.
If it walks like a duck, walks like a duck, swims like a duck, it probably is a duck.
So if it answers like someone intelligent, um, yeah, has some reasoning and, um, behaves intelligent, what do you miss with AI systems?
Where do you draw the line?
What’s, uh, I don’t know that I have any special, anything special to add.
I mean, we’ve been asking this question since Marvin Minsky in the 1950s.
What does it mean to be a human?
And, and, uh, and the Turing test, of course, even earlier, even earlier than that, I, I don’t want to say it doesn’t matter, but for the purposes of this conversation, it kind of doesn’t matter if that makes any sense.
Right?
So like using LLMs as, as tools in software engineering, they are very helpful.
There’s just no argument about that.
They can do a lot of things and we should feel it’s a new tool, just like a compiler.
Like there’s now things that we used to have to do manually and now we don’t have to, and we can let the machine help us with it.
That compilers didn’t make engineering go away.
Some people thought it would, by the way.
Um, uh, and COBOL would make things go away because we don’t need engineers.
You just have business people talk to it.
And we learned that wasn’t true.
Um, so I don’t know that it, I mean, it’s a super interesting question as a human and like from a philosophical, philosophical perspective.
Um, again, in the particular software engineering area, uh, I’m fine that it behaves.
I’m fine that it does what the tool does.
Um, as a, what’s the way to say this as a human?
I mean, I don’t want to be supplanted by a machine.
I don’t think anybody, um, and again, I don’t have any special insight into this to be honest, but I think the history of, I always want to look back and see what our analogies, the history of computing teaches us that we really don’t understand humans very well.
And every time we think, oh, well, that’s going to get the, get rid of the humans.
Like, nope, we just need more.
There are, there are more, more need for engineers.
So I think it’s just another one of, uh, one of those iterations.
I don’t know.
I’m, um, I’m, I’m not, I’m not concerned yet.
I’m concerned about many things about AI resource usage, the planet, all the bias that goes into the training, all the stolen stuff that’s gone into the, I mean, like there are lots of things to be concerned about.
What I’m not really today yet concerned about is it supplants human brains.
Yeah, quite interesting.
Um, I also wouldn’t compare it to, to humans because humans are even more than just the intelligence, but I think it’s quite close to, I mean, it’s a different form of intelligence, but quite close to the definition of intelligence we had during the last year.
Yeah.
And I think what we’re learning is our definition of intelligence, our definition of what it means to be human is evolving.
Every time we take, every time we figure out how to automate or industrialize a part of what it means to be human, we find new things that are unique about humans.
I don’t know.
A last question before we end this.
Um, do you think AI and LLM can be creative?
Sure.
Yes.
No, yes.
And, and slow.
Well, no, no, I, I don’t know.
Uh, I maybe I’m slow.
Cause I’m a human.
Um, the, uh, yes, if only because combining new combining existing things in new ways is creation.
How do you know that you can, uh, there was a time in my life where I did patent law briefly.
We will not go into that time of my life except for later, uh, some other time, but it is total patents are about what’s invention.
And you can have a lot of arguments about whether the definitions there are good and I have them, but combining idea, an idea from one discipline with an idea of another discipline in a new way, that is creation.
That’s how humans do it.
Um, so I, yes, uh, I, I think LLMs can be creative.
Can they do every bit of creation?
Again, this is philosophy that I don’t have any, I’m not any smarter, better at this than anybody else.
I have yet to see poetry or songs or music like the L and produced ones are just terrible.
I mean, imitate when they, they are getting better producing something that mimics a human that’s getting really good, but producing net new, uh, artistic creations, maybe I’m making me, I’m making a distinction between creative and artistic.
How about that?
Absolutely creative yet to see whether it’s artistic.
I don’t know.
Those are my sounds good.
So we are already quite over time.
Thank you for your time, for your insights.
Um, and I’m really looking forward to meet you, uh, for your, um, you have a keynote, it’s a software architecture gathering.
Um, so it’s, uh, the title is something like you can’t rewrite at all, um, about modernization.
Yeah.
So it’s a lot of the things we talked about.
So very briefly, the keynote, the keynote was inspired by the so-called SAS apocalypse.
So the software as a service companies for several months at the beginning of this year, I believe, yeah, this year, uh, something like $300 billion of enterprise value was removed from Salesforce, SAP, blah, blah, blah, because everybody thought nobody needs them anymore.
We’re all going to rewrite those soft, those pieces of software ourselves.
Anybody who really worked on them knows how insane that was.
And we’re seeing that the market has recovered, but I think that taught us a lot.
And we talked about a lot of these ideas here where the power of LLMs is teaching us modularity still matters.
Evaluation still matter.
Like invariance and architecture still matter.
In fact, they matter even more.
And so you can’t write, rewrite at all is my idea of, it might seem like I could just regenerate a customer relationship management system or a database or an airline reservation system by myself with my trusty LLM.
I don’t think that’s true.
Uh, and the, and the talk explores why that’s true and also, uh, the ways forward, if that makes sense.
Yeah.
So, um, yeah, I’m really looking forward to it.
I already, uh, blocked the conference slot so that I, I will be able to attend.
Thanks again for your time.
And, uh, thanks to all, um, yeah, or watching this stream.
Um, yeah.
Have a nice weekend.
Bye.
Bye.
Hi, I am Alaa Reuschenbach.
Do you organize any user groups, conferences, or other tech events?
Then feel free to add them to treff.tech, an uncommercial platform for tech events in the German speaking community.
It’s free without any advertising or tracking.
Just visit treff.tech or scan the QR code.
You will also find that link in the video description.
And by the way, you can find all software architecture stream events also on treff.tech.
And if you have any questions, feel free to reach out to me.