We’ve had a very recent uptick in engineers submitting PRs of hundreds of lines across multiple files, for Jira tickets that only asked for a one-line change. The engineers involved have been using AI assistants for nearly two years now, but there seems to have been a change in the last month or so in how aggressive the new models are at changing code.
Almost as if they’re paid by the token…
This is exacty what I’m seeing. I’ve been using it in a DevOps capacity to act on runbooks. The same wrote task 3 months ago now consume 4-5x more tokens. this correlated closely with when anthropic released auto mode.
Use ponytail to keep 50 line changes to 1 line, and use rtk to save tokens.
Also use an assistant to checkout someone else’s diff and review it in chunks
2 years and they still don’t know how to prompt…
I mean, I don’t know how to prompt. Mostly because I don’t use AI to code.
Yep.
We have one PR still blocked. Last change is a simple comment from me “Why ?”
The most important question that every change must answer.
Merge that shit, watch it all collapse, enjoy your forever holiday
“So, Daywim. Why did you let this obviously aweful PR pass your desk causing so much trouble for our company? I’m afraid we have to let you go because of this questionable performance.” - Corporate
What part of forever holiday did you miss?
I interpreted it as “Holiday that lasts forever because the company can’t work anymore” but I guess it is meant to mean “Holiday that lasts forever cause you got fired”?
Yes it was meant to mean fired
“Looks like I overlooked something in this 6k PR full of im meaningless dribble. Why don’t you ask the person who comitted the code how he overlooked this bug. Its his respinsibility”
Just throw the the slop creator under the bus.
Throw them under the bus by rejecting their PR. Integrity is your responsibility, the gesture is theirs. They’ll get shit for not getting their stuff done.
If you’re the reviewer you share responsibility if there’s an issue with the PR. Hopefully your teams culture is such that issues like that are treated as a learning experience, rather than a reason to pile on the individuals involved.
If you have someone on your team that creates 6k PRs with AI then you are in a loose-loose situation anyways. If you do a thorough review you will be done just in time for the next one and you will get in trouble for not getting your own shit done. So gambling on the PR is probably your best shot at survival.
Just don’t approve it? If they complain, they don’t have a leg to stand on.
“why don’t you have an AI checking that review?” That’s the reply I have heard from a big tech company thread. They will force you to be the one who take the fall when AI fail but will also shit on you for not using AI. That’s the whole point of the system.
If it can’t be sensibly be reviewed, reject it for that reason. Have them, or their agent, break it down into smaller, self contained, PRs.
If the company allowed people to use AI slop what makes you think they would have a culture to push back again that. Working in the bay right now and the culture have always been “move fast break a lot”
Our PR checks auto reject the PR if it has 1k changes
is it auto reject, or just doesn’t auto approve and leaves it open for manual review
It seems weird that you can’t do a pr at all with 1000 line changes, any moderate size feature addition could hit that mark
You know, you could just chunk it up in a way to keep it readable.
for a new feature request? a PR isn’t a commit, it’s a set of commits which would add to the line change amount.
Like even if you spread it out across 20 or 30 commits that’s still going to be the same line count.
I guess you could push not yet functional or used code to lessen the line count change, but that seems in bad taste. I’ve always gone off the working repo should always be in build or clean state and a push or commit shouldn’t break that.
I work in a Scrum team and we implement features iteratively, so we start with the minimal feature, merge it, get feedback, and go on from there.
At work there’s no excuse to keep a feature in a stale branch until you accumulate 1000 lines of code change.
Yeah, 1k is kinda a small limit but I get the logic, you can almost always break changes into smaller increments and not mass merge a mega PR that’s hard to review
Stupid question, but what happens to a rejected PR? Because features get built for a reason (there’s usually a Ticket/Story for the feature that the PR adds) so do those just get closed? Or does the person have to re-write the code entirely?
I know in my team, the most I could do is tell the coworker to self-review while keeping the PR open until they change some stuff. I have never rejected a PR before because no matter how bad a PR is, it always is technically necessary for the feature.
If a feature request requires changes of such a magnitude it is important to break them down into smaller chunks that can be reviewed either independently or sequentially.
I have never rejected a PR before because no matter how bad a PR is, it always is technically necessary for the feature.
Define bad? If the PR contains lots of unnecessary changes such as formatting or renaming simply tell the person to roll them back and come again unless they have very good reason to do so.
If the code quality is bad, well, that’s why you are doing the review. If all that matters was “Does it do what it is supposed to do most of the time?” some simple unit tests would be enough. Reviewing code means making sure it does what is supposed to do and does so in an acceptable manner. Criteria can be amongst others speed, security, ease of use or maintainability.
- Insecure handling of inputs? Add sanitisation and resubmit.
- Overusage of resource intensive features such as database queries? Group and optimize queries and resubmit.
- The codes formatting is not according to the internal style guide? Configure your damn linter and resubmit.
- …
If you don’t want to close PRs outright you can request new commits that fix the issues you identified, reevaluate the PR and decide again.
What if you remove some files? 🙃
We now have the ability to let Copilot review a PR on Azure DevOps, if someone sends a PR by Copilot I send Copilot right back at it
Guss Fring used to be a scientist? That a parallel to Walter White! Bravo vince, you’ve done it again
No Mr Bond, i expect you to approve
If LLM can make big PR, LLM can split PRs
The goal of AI providers is to make humans unable to maintain code, so you have to rely on their expensive subscriptions and tokens.
I’m not sure if you’re suggesting to use LLMs to review bad PRs, or that the author should redo them. Latter, for sure. Former, sounds horrible.
LBTM, rejected instantly.
Don’t worry half of those will be useless code comments
// Here I'm not using that other thing that is now completely irrelevant, but I'll leave a comment to the non-existing thing anyway because I'm avoiding it.REEEEEEEEEEEEEEEEEEEEEEEEEEEEEEE
The comment:
# This code does exactly what you asked: Never change state, only fetch the state and return the difference
the code: hallucinated database table drops
This one I’ve not seen. It seems fairly ok at not going completely batshit like that. But the design, layout and comments tend to be awful. It’s pretty good at solving isolated tasks though.
It seems you have a friend named Claude
Repeat after me: “Rejected. Reason: too large of a change for one PR.”
I struggle to review a 1k line change. When people give me such big changes I normally don’t believe they’ve reviewed them either.
That’s because they haven’t.
Mystery solved!
Try working on a codebase that’s all event-driven hexagonal CQRS with hand-crafted SQL for persistence. Add additional buzzwordy methodologies to taste.
Adding a single property to your product means you now have to update an aggregate class, several DTOs, and several event classes and handlers before you can even think about touching the UI.
And that’s in your main solution. There’s also at least one facade service you’ll need to make compatible and you also need to update the event simulator used for testing. The latter night involve having to touch every single line in a 2000 lines long SQL script.
Having to go though three separate 600-2000 LOC PRs for one PBI isn’t that exotic.
forgive them lord, for they know not what they do
Lord, I wish they knew what they do
deleted by creator
Occasionally I do that…but only because 500 of those lines are my comments explaining everything.
Please don’t explain that much, make your code easier to understand.
I’m aware. Not always feasible. For example, had to add a custom video capture solution that captures the last 30 seconds of a process for crash handling purposes.
You most definitely need to do add that much comments explaining the mp4 box format along with the box “hierarchy” of what is being written. Add to that MFT (h264 encode) code…
Basically, if anything, the comments are for me for when I look back at that code.
had to add a custom video capture solution that captures the last 30 seconds of a process for crash handling purposes.
Commanded from above? This smells like a noob idea.
I have no idea what your response even means/is getting at.
This is my attempt to translate: You had to add a feature like that ? Must have been an order from your boss. That seems like a feature someone rather inexperienced and unknowledgeable would request.
- Legal has issues with using existing libraries (including MIT). Definitely idiotic, but I can’t control that.
- Idea was mine.
- Subprocess that captures parent process active video with a rolling buffer for crash handling purposes is absolutely necessary when trying to reproduce issues in development (Gamedev editor). If you think associating the last 15s prior to a crash with a mindump isn’t helpful for debugging, I don’t know what to tell you.
The problem with Claude is that it doesn’t write code to be modular & reusable. Every tiny change requires a complete rewrite.
I’ve completely banned any code that can’t be explained. I’ve had my CTO send me code at 3 AM to implement and when I ask him what I’m looking at he just says it doesn’t need review, just push it.
Uhh, no sir, I’m not doing shit because you’ve handed me GCC and we’re MSVC.
After I bitched endlessly to the CEO about that he said I have final say on what goes into the project.
we’re msvc
Then switch to a real compiler on a real operating system, duh
deleted by creator
I’ve had my CTO send me code at 3 AM
I hope you don’t even respond until your next normal working hours!
I love my job, even when I have to deal with nonsense like that and I’m compensated very well to be on call 24/7.
No amount of money would make me put my health at risk like that. Been there at a job before where I was always working. No thanks.
that sets a really bad example. you shouldn’t do that to yourself and you shouldn’t allow it to happen to anyone else.
I’m going to retire at 40, I think I’ll keep going instead of listening to you.
This is where “that crazy old customer” comes from.
He knows best what’s best for him. If he’s explicitly paid extra to be on call 24/7, and he’s happy with that extra. Let him be om-call 24/7.
There are situations and jobs where 24/7 availability is needed. Someone has to do it. And if that someone believes he’s getting enough of a compensation for it, there’s nothing wrong with it.
Yup, I was on call 24/7… compensated very well and now I’m retired at 43. So, yea I’m going to have to go with letting the dude make his own choice as far as what’s best for him.
well, that depends entirely on your prompt/process, it can be done
Just more prompting will do + make no mistake
unironically, a better prompt does yield better results - shocking, I know
Unironically Claude (or whatever you use) will almost always deliver code that is shit, it’s just less shit if you prompt it better, LLMs are good to make short snippets if you get stuck tho; And remember to fucking check what the code is and rewrite bad shit
I’ve seen plenty of code in my life, from humans and AI.
For the past year or so, these agents can code just fine most of the time, as long as they are given enough context (or have the tools to get it). Regardless of how many downvotes I get here, they really are capable of generating decent code. I’m sorry you couldn’t make it work yet.
I’ve seen too as I have to review it. AI produces very professionally looking bad code.
It is so weird to explain, but that’s basically what I see.
Before LLM I could immediately tell someone’s code is crap, now I have to spend a lot of time trying to understand what it does to reach the same conclusion.
My honest opinion is that it’s bad because a lot of people using LLMs have no standards and push the first thing that seems to work. Be mad at who’s at the driving wheel, not the car.
You absolutely can generate crap with agents/LLMs, and like a humans writing, the first draft will probably be subpar or maybe complete garbage. Every new session is a clean slate, that’s why putting effort in the documents guiding it is so important.
But AI bad and human code is so awesome! Better downvote. But seriously, there are people here saying that LLMs can not add anything of value in any way. About as delusional as Republicans.
With enough seniour developer’s time and dedication you can spend days and enough water to flood a town, so you can badly maybe do something that a junior dev can do already (your shit will still be worse). If that’s not an achievement of a modern technology I don’t know what is.
It can but I’m not looking to make things even more complicated. We have enough unexplained non-sense in the project as is, having to link it correctly is just endless pain that I don’t want to deal with.
No
You ask your LLM of choice to look it over, completing the shit-cycle
I fully expect this to become the new normal being pushed by management.
“We identified PR reviews to be blocking our newfound AI-powered efficiency, so we are now mandating all the reviews to done by AI. Also we figured all the developers are now useless since all you do is ask Claude to solve tickets, so you are all fired”
I wonder how long it takes for the first high profile disaster happening because of a policy like that.
How is bun doing btw after their “We used 60 agents and 200k in tokens to rewrite in Rust” ?
Actually I think it’s doing well? The language to language rewrite is actually a strong suit of LLMs, as long as there is extensive years worth of tests to check rewrite behaviours.
I don’t have practical use cases for that strong suit though.
Oh, and bun seems to be dead now with 2.5k open PRs.
And merge to main takes over an hour.
And the total rewrite cost was significantly higher than the headline (multiple Prs by Anthropic employees, rough count 20% of total LoC of rewrite)
And there is still no release in sight on Github - but Claude is shipping with rust bun afaik.
Uh oh, new workaround against copyleft on zhe horizon?
It was higher sure, but anthropic also has some of the most expensive models out there. That cost could be 5x-10x less just by going with cheaper model providers (if one were to pay the API costs, not the case for bun).
bun 1.4 was released 3 weeks ago, btw
That cost could be 5x-10x less just by
My point was that it was cost of model rewrite by agents + 3 months worth of coding by lots of people.
5k open PRs atm + 3.5k open issues. I have no idea what is the state of Bun right now, but I am not confident in it.
I don’t think there was a lot of people working on the rewrite. Most PRs are from bots. Original estimates for a manual rewrite were a small team working for a year or so, which puts total costs over $1M. Even doubling the token cost estimates, it was still cheaper than doing it manually by a factor of 2x-3x
This is literally how corpo I work for wants us to work. They call it… Outcome baded review. But no bugs on prod lol.
This is the exact thing that majority of “AI” companies are doing. Shit in, shit out, nobody knows why or how things work
LGTM
Let’s Go Topple the Monarchy!
Let’s Gamble, Try Merging is my favourite
Dung ahead, try madness

















