L01 Deep Dive: Course intro + the AI coding landscape EECS AASE Podcast, season 1, episode 4 Transcript of the published audio. The episode is AI-generated from course materials; this transcript is produced locally from the audio. Two synthetic hosts in conversation. Turns are separated by blank lines and are not attributed, because the local transcription does not label speakers and guessing at labels would be worse than none. Okay, let's unpack this. Imagine you are showing up on day one of this really high level software engineering class. Right. You're probably pretty tense. Exactly. You've got your laptop open. You are bracing yourself for this dense wall of architectural diagrams and terminal commands. And the very first instruction you get from the front of the room is smile for five seconds. Which is just I mean, it completely derails your expectation. It really does. But, you know, welcome to our deep dive into lecture one of AASE, which is applied agentic software engineering, also known as EECS 498. Yeah, EECS 498. It's a fascinating class. It is. And we are working from a really great stack of sources today. We have the official lecture overview, the slide deck, and a completely de- identified transcript of the lecture exactly as it was actually delivered. So you're getting the real on the ground experience. Right. Think of this as your ultimate post-lecture study companion. Our mission today is to, well, extract the core philosophy of A.I. programming being taught here, understand that underlying mindset, and map out the exact practical steps you need to take this week. Because you really are embarking on a fundamentally new way of thinking about software engineering. Yeah, completely new. So let's start with that bizarre opening instruction. I mean, why spend the first five seconds of a highly technical class just smiling? Well, the instructor grounded this in the sort of peculiar piece of internet culture. There is this YouTuber named Ben who runs a channel called Sitting and Smiling. Wait, really? Just sitting and smiling? Literally. He just sits in front of a camera and smiles for four hours straight. Wow. That's intense. It is. But the point is the physiological act of smiling actually forces a shift in your body chemistry and your outlook. Right. And in the context of lecture one of A2SC, that matters immensely because handing over control of your code to an A.I. agent, it requires an incredible amount of psychological confidence. You have to be comfortable being in this unfamiliar kind of uncomfortable territory, which brings up a very real issue in engineering, right? Yeah. Imposter syndrome. Oh, absolutely. Because handing the keyboard to a machine can make you feel like, you know, you aren't a real programmer anymore. And the lecture tackled this head on. Yeah, they didn't shy away from it at all. The instructor acknowledged that even award winning professors and MIT grads, they feel completely out of their depth sometimes. Yeah, they actually had the room do this affirmation exercise. Everyone had to look at their neighbors and say, you belong here. Which is great because it establishes a baseline of psychological safety before the heavy technical work even begins. And that technical work is structured as this really slow, deliberate build. I mean, it's a 15 week arc centered around one single object, which is a coding agent. Right. And the course breaks this progression down into three very distinct verbs. OK, what's the first one? First is apply. And that happens in weeks one through three. That's where you use an existing agent to build one. Got it. Then comes analyze, which is weeks four through seven. And finally, create in weeks eight through 15. I kind of want to pause on that middle phase. Like, what does analyze actually mean in a practical sense? So that is where you physically take the human meaning you completely out of the loop. Wait, so you aren't doing anything? Well, instead of sitting there, you know, manually accepting or rejecting the AI's suggestions line by line, you give the agent a set of tools. You establish these strict stop conditions and you just let it run autonomously. Wow. Yeah, the goal during the analyze phase is strictly to measure what breaks. You are looking at the logs, analyzing the failure states and really understanding the limitations of the model when it doesn't have a human holding its hand. OK, so you are intentionally letting it fail in this controlled sandbox just to map its blind spots. That's wild. And then the create phase in the back half of the semester. Right. So once you know exactly where it fails, weeks eight through 15 are all about building robust automated systems to actually catch those failures. Makes sense. You grow that basic system into a full blown persistent assistant with, you know, memory and custom methods. Honestly, the wildest part of this structure to me is that the entire 15 week arc happens inside one single Git repository. Yes. Like you aren't doing these weekly throwaway assignments where you build some toy app and just forget about it. Your week to commit in your week 15 demo are the exact same repo. It's a very intentional pedagogical choice. You are basically living with your own technical debt. Yeah. Reminds me of like building a classic car out of spare parts in your garage. Oh, that's a good analogy. Right. Because someone else might just go buy a Corvette off the lot and drive it really fast. But if it breaks down, they have no idea what to do. Exactly. By building it yourself, you know, piece by piece over 15 weeks, you know exactly how the engine works, where the wiring goes and how to fix it when it stalls out on the highway. That captures the ethos perfectly. You start by just driving the tool to see how it feels. Then you step back to see where the engine knocks. And finally, you rebuild the transmission yourself. So with an arc that comprehensive, how are they actually measuring success? I mean, the syllabus explicitly states there are zero exams. Right. No exams at all. So how do they know if I'm actually learning to be a good curator or if I'm just kind of coasting? Well, they weight the grading heavily toward the outcome of those three phases we talked about. Ten percent of your grade is just administrative, mostly attendance. Which they track via poll everywhere, right? Yeah, they do. Which, according to the transcript, completely crashed during this very first lecture. It did. It was this fantastic, totally unintentional lesson in rolling with the punches when technology inevitably fails. Everyone on the roster ended up getting credit anyway. That's hilarious. So what about the other 90 percent? It's heavily backloaded. You've got 20 percent for apply, 25 percent for analyze and 55 percent for create. Ah, OK. So the stakes get higher as you take on more of that architectural responsibility. The syllabus also mentions these evening hackathons. And I usually associate those with like sleep deprived weekends fueled by energy drinks. But these sound a bit different. Yeah, they are highly focused, fixed task evening sessions. The entire cohort is basically in one room. And the goal is to pressure test your skills live in a really collaborative environment. And the first one's coming up fast, right? Thursday, September 24 at the ETF. Yes, the engineering facility. Perfect. And I think we should highlight perhaps the most remarkable administrative detail of all this. The course costs absolutely nothing beyond tuition. Nothing. There is no textbook to buy. And critically, there is zero L.M. spend required. You run local models on your own machine or use a cane backup if your local hardware can't handle the compute. And that emphasis on keeping it locally run and financially accessible. It ties directly into the core philosophy of the course. Because to successfully build that agent over 15 weeks, you have fundamentally changed your identity. You are no longer a typist. You have to become a code curator. OK, let's define the functional difference there. What actually separates a typist from a curator? So a typist focuses almost entirely on the how. They spend all their cognitive energy writing boilerplate, managing formatting and, you know, fighting with syntax errors. Right, the tedious stuff. Exactly. When you're a typist, owning the code means you literally generated the keystrokes yourself. But a code curator operates at a much, much higher level of abstraction. OK. A curator specifies the what and the why. You describe the overarching intent, you allow the machine to generate the initial draft, and then your primary job is to rigorously evaluate that output. So you review, you redirect, and then you accept. Yes. And this ties directly into the course's integrity rule, right? If you can explain it, it is yours. That's the golden rule. It totally shifts the definition of correctness. Correctness no longer means, like, I typed this perfectly. It means this works. And I can defend every single line of this architecture, regardless of whether an LM typed it or not. Exactly. You own the outcome. Here's where it gets really interesting, though. Yeah. Because I have to push back on this a little. OK, let's hear it. If the AI is writing all the syntax, and my job is merely to review and defend it, why did I just spend two years suffering through EECS 281? Ah, yes. Like, why did I agonize over all these complex data structures and sorting algorithms if I'm just going to ask an LM to write my functions? Doesn't this kind of make traditional foundational computer science obsolete? Well, it feels that way right up until your production server catches fire. Oh, man. The lecture actually addresses this specific anxiety with this brilliant metaphor. You do not give a child a nail gun. You give them a little plastic hammer to hit plastic pegs so they can understand the basic mechanics of swinging, striking, and force. Right, so they learn how the tool interacts with the material. Exactly. AI is the power tool. It is a commercial-grade nail gun. You absolutely need the rigorous foundation of EECS 281 so you know exactly where the nail is supposed to go, what kind of wood you're firing into, and what structural load it can bear before you pull the trigger. Because if you just hand a nail gun to someone who doesn't understand plumbing, they will confidently drive a nail right through a high-pressure water pipe. Precisely. AI has vast encyclopedic knowledge, but it has zero judgment and zero taste. Zero taste, yeah. The lecture highlighted a highly specific architectural disaster, just to illustrate this. Say you ask an AI to match some user records. It happily suggests an N-squared lookup. Okay. It writes the code beautifully. It passes your local unit tests. It works flawlessly when 10 users are on your staging server. But N-squared complexity scales quadratically. Right. So the moment you push that to production and a million users hit your site, that N-squared lookup creates an exponential bottleneck and literally melts your server. Wow. And the AI will not warn you about that. It just executes what you asked. You need that foundational computer science knowledge to spot those subtle, delayed-fuse architectural disasters before they happen. So we are holding this incredibly powerful, potentially dangerous nail gun. How do we actually steer it? The lecture mapped out the historical landscape of AI coding tools to kind of explain our current position. Yeah. We have seen four distinct waves. Wave one was autocomplete around 2021 to 2022. Okay. Wave two was chat, basically 2022 to 2023. You ask a question, you get a block of code, you manually paste it in. Let's focus on where we are right now, though. Wave three, the egenic tools from 2023 to 2025. This isn't just some chat bot in a browser anymore. No, not at all. Tools in wave three actually read your local files, they edit across your entire project simultaneously, and they can even run your tests. Yes. Wave three tools have systemic visibility, and they are really the stepping stone to wave four, which is orchestration from 2025 onwards. And what does orchestration look like? Wave four involves long-running supervised agents that maintain state, use external tools completely autonomously, and actually possess memory. To bring back that vehicle metaphor, wave one autocomplete is like driving a manual transmission car. Wave three, where we are right now, is like typing an address into a GPS while you still have your hands on the steering wheel. Yeah. You are navigating, but you are heavily guided. Yes. Wave four orchestration is like sitting in the back of an autonomous taxi. You just provide the final destination, and the agent handles the traffic, the route, and the pedals. That is a highly accurate way to visualize the progression, and I think it cuts through a lot of the fear-mongering out there. Yeah. The lecture makes a crucial point about the reality versus the hype here. AI is not replacing programmers. Definitely not. Engineers who know how to expertly use AI are replacing engineers who refuse to adapt. Exactly. And to be the engineer who adapts, the lecture says you have to master the big three levers. These are basically your primary control dials, context, model, and prompt. Let's break down the mechanics of those, starting with context. So context is the information and environment you create for the AI before it even starts generating anything, like which files does it have access to, what documentation is currently visible to it. Got it. And the model. Model is the specific neural network engine you choose to run. That dictates the balance between raw reasoning capability, speed, and computational cost. And finally, prompt. Right. Prompt is the actual linguistic instruction you provide to specify your intent. And when it comes to that third lever, the prompt, the golden rule they emphasized in the lecture, is the KISS principle. Keep it incredibly simple and focused. Keep it simple, stupid. Yes. The instructor gave this great example of extracting even numbers from a list. The instinct for a lot of beginners is to just over-explain everything. Oh, absolutely. They write this rambling paragraph like, could you help me write a Python function that iterates through a list, checks for evens, maybe use a list comprehension because I think it's cleaner, but also make sure you handle empty inputs. And when you prompt like that, you are just muddying the waters. Right. The model's attention mechanism gets totally scattered across conflicting instructions, hypothetical edge cases, and all your personal preferences. A tight, focused prompt almost always yields better architecture. So what's the superior version? Simply, write a function that extracts the even numbers from a list. It is exactly like micromanaging, a highly capable but intensely literal- minded employee. Yes. If you give them a clear, unified goal, they will execute it with precision. But if you stand over their shoulder and give them five contradictory caveats, demand a specific methodology, and throw in a hypothetical edge case all in the same breath, they're going to freeze up and make a mess. Exactly. You have to give the model the space to infer the optimal mechanics. This transitions perfectly into the actual tooling we're using to put this into practice. Yeah. But I'm genuinely stuck on the software choice here. Why is that? Well, if our goal is to be high-level orchestrators, why are we being forced into a bare-bones command line tool like Adr? Why not use a beautifully polished GUI like Claude code or cursor? Well, it is a very deliberate pedagogical constraint. Highly polished tools often abstract away the friction. Yeah, they automatically manage your context window, and they hide all the prompt engineering behind these slick interfaces. Adr is used specifically in the apply phase because it forces you to confront those big three levers nakedly. There is no graphical interface hiding your context decisions. Exactly. If you want the model to see a file, you have to explicitly add it to the chat yourself. It forces visibility. Plus, Adr uses an open-source, OpenAI-compatible API, right? It does. That means swapping your model is just changing a single line of text in your configuration file. You aren't locked into a specific vendor ecosystem, which is how the course ensures you can run it for free on local hardware. And starting in week three, Adr is structurally small enough that you're actually going to tear it apart and rebuild its core logic yourself. That is so cool. And to demonstrate how effective this stripped-down environment can be, the instructor actually performed a live pair coding demo. Yeah, that was impressive. They connected Adr to a local 9B model. That's a model with nine billion parameters and executed three coordinated changes completely simultaneously. Right? They asked the agent to refactor a Parseregs function to handle empty arguments. They requested a unit test specifically for that new path. And they had it replace a clunky if chain with a clean dictionary lookup. And what happened? All the tests passed green and the instructor typed absolutely zero characters Python. And because they were monitoring the token usage directly in the terminal, you could see the entire interaction cost pennies. It completely demystifies the financial barrier of AI engineering. Which brings us to your immediate deliverables for the class. You should have already completed the Adr setup in lab zero. Right. That was step one. This week, you are diving into the Tasker repository. When you accept the GitHub invite, you'll find a repository structured around 16 guided lessons based entirely on failing test cases. For part one, you are working through lessons one through 10, building out a command line application feature by feature, driving the AI until the tests pass. But there is a massive catch in the instructions. What is it? They explicitly demand that you start the Tasker lessons using a weak 4B local model. Okay. Yeah. If we are trying to become hyper efficient curators, why would we intentionally hobble ourselves with a weak 4B model when a much smarter 9B model is sitting right there? This raises an important question about how models process information. A 4B model has fewer parameters, which means its ability to infer contextual leaps is severely limited. It's a fragile tool. Right. If you give a 4B model a sloppy context window or a vague prompt, it will fail catastrophically and instantly. Just breaks. Exactly. It does not have the sheer processing power to guess what you meant or to silently patch over your logical errors. It forces you to learn rigorous prompt discipline under immense pressure. I see. The frontier models actually enable bad habits. They do. If you use a massive model, you can write a terrible muddy prompt. And the model's massive parameter count allows it to guess what you want anyway. Yes. But you haven't learned how to steer the machine. You have just learned how to be a passenger. The weak model exposes your flaws as a curator. Precisely. The 9B model is only recommended for part two next week once you have felt the pain of wrestling with the 4B model. That is pretty clever. The curriculum really wants you to experience exactly what a larger parameter count buys you in terms of architectural reasoning, rather than just reading a benchmark chart. Developing that firsthand intuitive judgment is exactly what you will be graded on in the analyze phase. So what does this all mean for you this week? Let's recap the critical takeaways. Sure. You are fundamentally shifting your identity from a typist who worries about syntax to a code curator who owns intent and outcome. That's the big one. To get that correct outcome, you must master the big three context model and prompt always keeping the KISS principle at the front of your mind. Keep it simple. Your immediate actionable steps are to verify you have passed the lab zero setup gate and to complete part one of the Tasker lessons using that unforgiving 4B model. All of this is due Friday, September 11. And as you write your first prompts this week and you watch that 4B model inevitably struggle with your instructions, I want to leave you with a broader long-term question to mull over. Okay, let's hear it. If the day-to-day reality of being a software engineer is shifting away from memorizing syntax and moving almost entirely toward curating AI outputs, will the ability to seamlessly integrate, direct, and negotiate with autonomous AI agents soon become far more valuable on a resume than mastering any specific traditional programming language? That is a massive paradigm shift. It really redefines the entire value proposition of an engineering education. It absolutely does. So as you sit down to tackle those failing tests in the Tasker repo today, remember that weird instruction from the very beginning of the lecture. Give yourself a five-second smile. It helps, I promise. You are navigating unfamiliar waters. The tools are definitely going to break. And you are going to have to rely on your foundational knowledge to put the pieces back together. But you're mastering a power tool that will change how you build software forever. Grab your plastic hammer and we'll see you in the next deep dive.