Two years ago, when I first started experimenting with agents, what I built was mostly a prototype for a limited set of use cases.
A lot has changed since then.
As LLMs have become more capable, our Data Platform Agent has evolved into something that is now widely used across the team—for SQL-related questions, on-call alert triage, lag monitoring, and scheduled workflows for fixing bad data.
Building and evolving it over the past two years has also changed how I think about agent harnesses.
Harness may diminish as model get stronger: Much of the progress in agent capabilities comes from improvements in the underlying llm—their ability to reason multi-turns, use tools, and operate over longer context. A harness does not create that intelligence; its role is to expose and organize those capabilities for a particular application.
Model can overfit, same for harness: Different environments require different designs: a coding agent, a research agent, and a conversational agent may need very different control structures. Even within the same domain, a newer or more sophisticated architecture is not automatically better. Evaluations, instead, should guide architecture decisions rather than architectural novelty itself.
Does it still make sense to talk about harnesses? I think it does. Research directions such as recursive self-improvement (RSI) and test-time-training (TTT) adaptation point toward agents that can increasingly improve their own behavior and updating their weights during runtime. But we are still far from production agents that can reliably redesign their own operating environment end to end. In the near term, I expect harnesses to become more lightweight, but not disappear.
For me, this is what makes agent engineering so interesting. It feels a bit like building with Lego while new pieces are constantly being added to the box. While you are figuring out how to assemble the pieces you already have, new models, tools, protocols, and abstractions keep expanding what you can build. Some designs become obsolete quickly; others suddenly become possible. There is a strange feeling of being pushed forward by the pace of the field—but that is also what makes the journey exciting.
I see this as a problem of people losing grip on the trajectories as models get more capable. Two reasons: you either become more trusting of your agent, or you don't know what it's doing because the CLI is no good at presenting complex information.
Speaking of which, has anyone used a cost-visibility UI like AgentCost or Langfuse? (not affiliated, just curious)
> I think Neovim's decisions are justfiable in this context.
Needing to know in advance that neovim requires a specifying a different undodir to vim, because neovim's persistent undo format is unstable even though vim's is stable, but the feature is called the same thing and neovim bills itself as a plug and play replacement for vim...
This is very predictably going to lead to loss of user data.
It's impressive that something this simple can do this much. But those funky types of actuators all live and die by transfer learning now.
If a robot AI can figure out how to operate them with very little sim and teleop data, and learn to take advantage of their strengths while maintaining good performance on tasks learned from UMI datasets, teleop data or human headcam videos? Allowing the same "robot mind" to work with different actuators?
Then I would expect those to have a decent niche - sitting between the classic two finger UMI gripper and a humanoid hand. Not a drop-in replacement for a human hand, but still more dexterity per hand without sacrificing all of the ruggedness and mechanical simplicity.
If transfer learning for different actuator types doesn't work so well? I expect the field to collapse to a binary of "UMI gripper or humanoid hand", with nearly no in-between.
In general, I'm carefully optimistic? But we are yet to demonstrate with confidence that this kind of transfer for actuators with radically different kinematics would work.
You've got it backwards. You've used AI-adjacent jargon to blindly grope towards philosophical issues that have been debated for thousands of years. You may enjoy reading actual philosophy, you might be surprised how deep they've gone on those issues. Thousands of years is a lot of time.
From a Christian perspective, I imagine a lot of what you've said is incoherent. God is all-knowing, so he's not 'worried' or going through 'crises'. He already knows how the story ends. It's doubtful he even experiences time.
The Certainty view is a nice touch. Is that edge band estimated from the source pixels, or by tracing a few slightly different versions of the image? I'm curious how it behaves on JPEG halos around a logo.
I will deliver a detailed plan for a graphical desktop for the ZX Spectrum. I will complete this within 7 days and deliver it via email to the provided address.
Article is all over the place. First it says X, then Y, then says both X and Y are valid. I mean sure, design-work isn’t always clear-cut, but this article can’t make up its mind about anything and just comes across as not even trying to convey any sort of information.
I'm not gonna gather 'evidence' for or against what some guy who happens to maintain a text editor I use thinks about the issues of the day. It's just creepy behavior, and a huge violation of personal boundaries. Why do you think this is an acceptable thing to do? Go outside.
Regardless, the material you've posted doesn't indicate much about him to me beyond that we have political differences, which is okay. That's literally fine. It's definitely not enough to condemn someone as a Nazi sympathizer.
I'm not arguing that it's good. But the average person doesn't write well either. Similarly, the average web app doesn't have good accessibility, probably doesn't render right on some phones, etc. We're not comparing to some Platonic ideal. AI content stands out mostly because of sameness.
The moon is 400 times smaller than the sun, but the sun is 400 times farther away, so they appear to be the same size. The solar eclipse we see, with the moon obscuring the sun but just allowing light around the edges the way it does, may be a very uncommon phenomena.
It started around 2015 when they went balls-deep into DEI nonsense. Prior to that Google was an engineering company. After that, it was a degeneracy. We still laugh at Google to this day, all the way to the bank.
The perennial question on HN being "is Rails dead?", I just somehow didn't expect DHH to be the one driving the stake in its heart.