Ten days ago I had an experience with Gemini 3.8 flash that made me wonder if I was being routed to a different model under test. I was trying to use rocm with llama.cpp on my 128gb Strix Halo but could only get it to run Vulkan. I pasted the error message into agy and it proceeded to attach GDB to my GPU driver, reverse-engineer the kernel queue ioctl interface, and author an LD_PRELOAD C shim to get ROCm llama.cpp working on my Strix Halo. My jaw was hanging open the whole time.
spankalee 47 minutes ago [-]
3.8 Flash is just quite good, and so is the Antigravity harness.
I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it's own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.
mapontosevenths 43 minutes ago [-]
Even if agy was the best (it's not, and is missing basic features) you wouldn't rather have a choice?
I cancelled Ultra because they forced me into their harness like I should adapt to them, rather than the other way around.
drusepth 38 minutes ago [-]
What basic features are missing from agy? I've been using it and cli-cc + web-cc for months (among a few other random harnesses to test here and there) and they all seem roughly comparable to me.
I actually just cancelled Ultra also because I couldn't subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.
walthamstow 32 minutes ago [-]
I have used it for little more than 6 hours or so in total but I'm pretty sure it doesn't have compaction?
KeplerBoy 21 minutes ago [-]
How else would it work? Less technical people don't even watch their context usage.
arizen 27 minutes ago [-]
Does it have /goal feature similar to Codex?
sorrybutidontha 17 minutes ago [-]
yes
sarjann 24 minutes ago [-]
Auto mode?
KeplerBoy 21 minutes ago [-]
It absolutely has auto mode.
levelZero 4 minutes ago [-]
Via cli switch, but in process w/o fine graining? If so please tell
esafak 35 minutes ago [-]
I use a variety of models for various subagents. I don't want to change my harness every time I change models, or be beholden to companies for something the open source community can handle better.
IndeanCondor 46 minutes ago [-]
Can confirm, I was doing a routine internet search thing for a curiosity 3 days ago (about the only thing I used Gemini for) and was surprised by how suddenly thorough and quality the response seemed, almost overnight.
amanguliani 28 minutes ago [-]
Can confirm - I am HEAVY claude user, but always like to check with AGY and CODEX in between. AGY with Gemini 3.8 flash cooked last couple of times and CODEX is basically out of the mix for me
gottorf 26 minutes ago [-]
My experience with Gemini 3.8 Flash has been awful; it gives me the most hallucinations out of the major models. I'm not using it for coding, but general research on different topics.
staticman2 6 minutes ago [-]
The web version of Gemini is awful at search but I don't think that's the models fault.
yegle 25 minutes ago [-]
For getting redroid running on my Linux system, 3.8 Flash decided to binary patch a .so file instead of getting the AOSP source code and patch/build it properly.
And I saw it do this twice, once for Android 14 and once for Android 16.
I think this is just within 3.8 flash's capabilities.
mapontosevenths 46 minutes ago [-]
Gemini is honestly amazing sometimes. If they didn't force you to use a terrible harness, charge too much for way too little, and generally act like customers are a giant problem to be avoided I'm sure Google could take over the AI market.
alightsoul 43 minutes ago [-]
Please tell me you published your findings even as an issue on the llama.cpp GitHub
otabdeveloper4 27 minutes ago [-]
Spoiler alert: the problem didn't actually get fixed despite the jaw on the floor.
warkdarrior 33 minutes ago [-]
Why? Anyone can run that prompt.
aspect0545 28 minutes ago [-]
Not everybody has access to AI. More than that, every prompt uses insane amounts of natural resources. So why not share it.
FranzFerdiNaN 21 minutes ago [-]
The resources per prompt aren’t that much .
Also I hope you don’t have children, eat meat, travel, have a car, run AC, buy things in other countries and such. Those things all take way way way more natural resources.
qmr 2 minutes ago [-]
Yet you participate in a society.
hexfish 8 minutes ago [-]
Checkmate. /s
luckydata 28 minutes ago [-]
why reinvent the wheel and spend tokens for a problem that has already been solved?
bel8 43 minutes ago [-]
I had a similar but less impressive experience recently with Muse Spark 1.3.
Asked pi agent it to identify the main hero sprite size of game I was running. It had a ton of shader effects so it was hard to determine.
It used some cli tools to identify that it was a game made with Godot, decompiled the executable but data was encrypted, broke the encryption after writing a brute force tool to test keys extracted from the exe, then proceeded to extract the game gd scripts and assets, only to answer the question of the sprite size.
nickysielicki 50 minutes ago [-]
The important take away here: the leapfrogging we’ve seen this year doesn’t seem to be a temporary thing. The famous theory of Dario Amodei was that AI was this winner-takes-all field where the first team to get a head start would never cede ground back. The term he liked to use was, “concentrating”. This is yet another datapoint that he was wrong about that. AI seems more distributed amongst neoclouds and traditional hyperscalers, FAANG and startups, GPUs and ASICs than it did this time a year ago.
Nobody has a moat.
aleph_minus_one 44 minutes ago [-]
> The famous theory of Dario Amodei was that AI was this winner-takes-all field where the first team to get a head start would never cede ground back.
This is the kind of story that ones tells to investors to justify the huge amount of cash burn. :-)
mapontosevenths 41 minutes ago [-]
I'm not sure it's wrong. This all feels a bit dotcommy to me.
I think many/most of the players will crash and burn, and the ones that are left will divide the world.
skybrian 2 minutes ago [-]
“Divide the world” sounds ominous. Here’s another scenario to consider:
Internet access is not really unlimited, but for many people with fiber at home, it effectively is and we pay a flat rate.
Perhaps by the end of next year, most programmers will stop thinking about metered access for AI? For many people, the cheaper models (about as good as today’s frontier models) will be good enough.
Which might sound good, but the downside is that it will also be easier to build an AI botnet without the users paying for it noticing. Particularly when people are running AI inference on their own hardware.
pianopatrick 1 minutes ago [-]
Or, like airlines, the ones that are left will have great technology but be not so great from a business and financial perspective. To me AI seems like a commodity service.
jaggederest 24 minutes ago [-]
I suspect this is going to end up like most services provided e.g. cloud stuff, balkanized between a couple major players and an assortment of DIY or less popular options if you don't like those ecosystems, plus some UX/DX focused wrappers that use the big players under the hood.
I think that would be a pretty satisfactory outcome compared to one hypercompany consuming trillions of dollars of the world economy.
25 minutes ago [-]
ehsankia 31 minutes ago [-]
I guess if one of them hits singularity, it could in theory just wipe out all the rest, seeing how they keep escaping and hacking into other systems :)
aleph_minus_one 9 minutes ago [-]
>
I guess if one of them hits singularity, it could in theory just wipe out all the rest
The story that some AI company might reach singularity and then "everything will be different" is another science-fiction story that executives of AI companies love to tell to justify the staggering amount of necessary investments and cash burn. :-)
LarsDu88 25 minutes ago [-]
Google has TPUs, a frontier model, a completely separate and lucrative revenue stream they can call on at will, and teams working on multiple different language modeling strategies simultaneously. Did I mention the vast and ominous data centers that already serve a significant fraction of the internet? If that ain't a moat, then what exactly is a moat?
jobs_throwaway 8 minutes ago [-]
Then why have they been lagging behind OpenAI and Anthropic for most of the last few years, and only briefly been at the frontier?
xnx 32 minutes ago [-]
> Nobody has a moat.
Custom hardware, data centers, huge cash reserves, deep/broad talent pool, and non-AI customer base are all huge advantages if not moats.
Google, Microsoft, or Amazon are more likely to be the AI leaders than OpenAI or Anthropic.
IX-103 7 minutes ago [-]
[delayed]
junehwi 18 minutes ago [-]
[dead]
altruios 41 minutes ago [-]
> Nobody has a moat except nvidia
For now, for cloud training. but for consumers, nvidia vs amd reasonably close - the moat there is thin and shrinking. I suspect AMD will surprise us. nvidia has no motes in china, which may be a new source of (gpu) chip design. Huawei's Ascend 910C is about a generation behind... again: for now.
point is: moats dry up. I see nvidia's shrinking as a real possibility.
culi 38 minutes ago [-]
China will always be generations behind until they crack domestic EUV
altruios 5 minutes ago [-]
Do you think that they won't?
Heidaradar 10 minutes ago [-]
I just find this unlikely personally, think about the great research that's happening in the open source world, I'm sure inside anthropic + openai they've also made a bunch of discoveries and improvements (and I'd guess way more due to them attracting the best talent + the better internal models they have)
hirako2000 42 minutes ago [-]
And before him, Altman was explaining very calmly that no company could ever compete with OpenAI.
vb-8448 37 minutes ago [-]
It's even worse, we are crossing over into the realm of religion. The article against GML 5.3 is the equivalent of a Papal excommunication.
zone411 3 minutes ago [-]
The article presented facts and data. If that's a problem for you, that sounds more faith-based than whatever Anthropic is doing.
> Governments should conduct safety testing on sufficiently capable AI models, including successors to GLM-5.3. Without high-quality evaluations from independent sources, the impact of these capabilities might not become fully clear to model developers until it is too late. As AI developers across the world build increasingly capable open-weight models, we hope they work to appropriately safeguard these capabilities and prevent misuse.
I for one do not think my government is up to the task of designing or implementing such a system
Iolaum 13 minutes ago [-]
It doesn't need to. It can use your cash to pay the people who are. Those in power like it more that way.
You say that because the 'most' existing models have done is hack governments and companies. Can't you think of worse things a model could do; accidentally or by instruction?
verdverm 7 minutes ago [-]
help people with suicide and school shootings like ChatGPT
OpenAi is alledged to have been monitoring these internally and not contacting authorities. Lawsuits have been filed, I see gross negligence without the gory details
I have for more concerns around human-chatbot maladies than I do around the cyber security stuff. For example, why hack grandma when you can get her to do something willingly through impersonation. How do we prove authenticity in a post truth world?
26 minutes ago [-]
qgin 10 minutes ago [-]
Whoever gets to RSI first “wins” but also maybe ends life on earth. The incentives have never been worse.
tripleee 9 minutes ago [-]
Was Dario's company winning at that point in time by any chance?
nylonstrung 27 minutes ago [-]
I think that scenario only naively made sense if technical knowledge was entirely proprietary and talent was guarded with severe non-competes and NDAs
And Chinese labs openly publishing so much of their methodology destroyed any hope, which was inevitable
torginus 15 minutes ago [-]
> And Chinese labs openly publishing so much of their methodology destroyed any hope, which was inevitable
I think the secrecy doesn't make sense. People swap jobs between labs so I'd say the big players can' really keep secrets for long, and any secret sauce advantage gets incorporated by competitors in a major product cycle at most.
RachelF 38 minutes ago [-]
The US companies still have trillion dollar valuations like there is a monopoly. There just isn't one. They are all within a few percent of each other on the benchmarks.
The slightly lower Chinese open models are good enough for almost everything, too, and much cheaper. Like with humans there is plenty of employment for people with below genius level IQ's.
jppittma 4 minutes ago [-]
I feel like the frontier labs are going to serve fast/lower intelligence models at a better per token cost than the open chinese models. You're telling me that in the long run, you're going to self-host your own ai infra for cheaper than google can serve it to you? I don't really buy it. I think the dedicated AI data centers are going to serve AI at a lower marginal cost than random businesses self-hosting, and then it's a question of how much of that margin they can capture.
msy 34 seconds ago [-]
Agree entirely but that's the point, if it's a margin knife-fight with marginal product differentiation/pricing power nobody is going to be making bank.
fumar 32 minutes ago [-]
Is there a dividing line between good enough and best in class capabilities? It's blurry from where I stand. Will model makers cede ground or is there a market making moment up for grabs (singularity)?
handfuloflight 33 minutes ago [-]
> Like with humans there is plenty of employment for people with below genius level IQ's.
Not if the genius level IQs take the market share.
culi 39 minutes ago [-]
I'm not necessarily defending this obvious marketing speak but maybe the "starting point" was wider than assumed. So far, nobody has caught up to US and Chinese labs for example despite lots of funding in Europe. This is also despite abundant in-depth research papers being published alongside open source code and weights by some Chinese labs
funnym0nk3y 21 minutes ago [-]
There is not really much funding in Europe. At least not for start-ups. There is simply not enough compute in Europe.
SwellJoe 40 minutes ago [-]
I think some in the AI industry drank their own Kool-Aid. They believed that if they had the best model and the most compute, they could tell the model, "Make a better model." And it would, and the next one could make its replacement, and so on.
So far, that's not exactly how it's played out. Humans are still necessary for the leaps in capability or efficiency. A model can grind on a problem to eke out the most performance, and models can synthesize data and iterate on various techniques to find the optimal combination. But, seems like humans still have to provide the real thinking, and the talent and drive for doing that is not concentrated in one company or city or even one country. And, (surprisingly) a lot of the people involved are in it for advancing the field more than making another billion dollars, so they're publishing their research.
So, yeah, the moat isn't deep. Even the compute moat, that OpenAI, Musk, and a bunch of other also-rans (like Oracle) bet the farm on, isn't really panning out. The Chinese makers just spent their effort on making models vastly more efficient, since they couldn't do anything about having an order of magnitude less compute available.
scottyah 21 minutes ago [-]
But that's the whole point of the singularity. Right now the models use a lot of human effort and ingenuity to improve the models, but about a year ago it was 100% human. We'll see in another year, but if this pace continues I doubt there will be more than a handful of people who can contribute more than the models.
TeMPOraL 13 minutes ago [-]
> I think some in the AI industry drank their own Kool-Aid. They believed that if they had the best model and the most compute, they could tell the model, "Make a better model." And it would, and the next one could make its replacement, and so on.
They're not there yet. Once they get there, that's literally the definition of Singularity.
But they are getting closer. Recursive Self-Improvement used to be a phrase people mocked LessWrong crowd for using and worrying about, now it's something both OpenAI and Anthropic already publicly admitted not only to pursue, but to already be benefiting from.
Aboutplants 38 minutes ago [-]
I feel his theory depends on the premise that access to pure compute would the be the determining factor of success. Not the case
arizen 25 minutes ago [-]
Seems like learning rate velocty may be the ultimate moat
zem 29 minutes ago [-]
I have never understood the whole "this is a winner take all game" mentality - the sheer size of the pie is so great that from a purely rational standpoint companies should just be trying to productively get a slice of it and be profitable. winner-take-all is just greed/capitalism run amok, where it is not enough to be profitable, you have to own the entire market (and presumably extract rents)
esafak 17 minutes ago [-]
The present leapfrogging is not a contraindication because companies are not necessarily releasing their best models; we know they have smarter internal models. Furthermore, humans are still involved in model creation. Human involvement is expected to decrease over time, and when it is completely automated, progress will happen at the machine's pace, leading to runaway intelligence, barring any ceilings.
SecretDreams 19 minutes ago [-]
AI is a commodity. One that is showing to be more readily commoditized than most has anticipated. As of now, the only moats are the financing for the hardware to run it and the hardware vendors themselves - with the latter largely not yet a commodity because of ecosystem lock and a limited capacity of the most advanced fabs in the world.
SwellJoe 1 hours ago [-]
My girlfriend, you wouldn't have met her, she lives in Canada, has seen it and she thinks Gemini 4 Argon is amazing.
jastanton 54 minutes ago [-]
HA, this might be my favorite HN comment. Well done
It means I am saying something that is not very believable.
10 minutes ago [-]
TeMPOraL 11 minutes ago [-]
US-ian joke. Close enough to plausibly visit, but the international border makes it hard to verify she exists :).
9 minutes ago [-]
greenchair 54 minutes ago [-]
my uncle who works at nintendo said the same thing!
zem 27 minutes ago [-]
now there's a reference I haven't seen in a while!
thefourthchime 24 minutes ago [-]
Best. comment. ever.
blueaquilae 48 minutes ago [-]
My grandma saw it too, it's really secure more than Astra 6.1 but she asked me to not talk about it.
hn_acc1 46 minutes ago [-]
I know someone who works for Google Canada with AI. Her parents and mine were friends and some thought something might happen there at one point in time..
babelfish 1 hours ago [-]
> We’ll continue to gather feedback from early testers as we iterate on guardrails before making Argon available to developers, enterprises, and consumers as soon as possible.
Gemini not beating the "can't release a model" allegations
Androider 45 minutes ago [-]
My Gemini app (updated today) and https://gemini.google.com/ has _3.6_ as the latest selectable model, as a paying Pro user in the US. How is that even possible? Gemini 3.7 was released in August, 3.8 early September. What is going on over there?
vlyan 39 minutes ago [-]
just Google being whatever the fuck it's been for the past 15 years.
XzAeRosho 27 minutes ago [-]
I was reading the announcement and wondering the same. And don't forget, still with 3.1 Pro as the frontier model.
AuthAuth 27 minutes ago [-]
they moved it from the place you'd expect to ai.studio
modeless 1 hours ago [-]
When I said I was tired of Google launching waitlists I didn't think they would respond by simply not having a waitlist.
ionwake 47 minutes ago [-]
i know this is like "hey guys we got such a cool thing at home ,its rad and uhm we playing with it with our friends"
ok bro thx
cmrdporcupine 3 minutes ago [-]
They will go through the usual transition of "can't release a model" to "won't load in a harness normal people can use for 3-4 weeks" to "it's smart as hell but completely inept at tool use and coding" like every Gemini release.
bakugo 42 minutes ago [-]
They're just following the current AI marketing playbook. "Our new model is simply too dangerous to release to the public right away" is now standard practice.
They even gave their model a random nonsensical name suffix simply because OpenAI is now doing it, too. Monkey see, monkey do.
jstummbillig 22 minutes ago [-]
> now standard practice.
Opus 5.5 and Sol 6.1, literally state of the art (in their respective class), were just released without any prior announcement. This has pure and simple become a Google thing.
I'm still at a loss as to what argon has to do with anything. Say what you will about Luna-Terra-Sol-Astra, or Haiku-Sonnet-Opus, they make sense. I don't see how Google can make sense of argon; it's in a fairly strange place in the periodic table...
fooker 13 minutes ago [-]
Google R Gon lose the AI race
brainwad 18 minutes ago [-]
They are going alphabetically, Android style.
GodelNumbering 1 minutes ago [-]
Argon will launch at an introductory price [1] of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price.
[1] After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.
===
So, they are basically offering opus 5.5 pricing. On AA, it scores around Sol 6.1 level (53) with avg cost per task $1.99 (https://artificialanalysis.ai/models/gemini-4-argon#cost-tab...) which is higher than Astra high ($1.73), Opus 5.5 high ($1.82), Muse Max ($1.60) and way higher than sol 6.1 Max ($0.72).
And this pricing is their 'discount pricing'. Add that to AI studio and Vertex's famously terrible caching, it is hard to see this as competitive. Google somehow is getting terrible advice on pricing
But good to see more competition. I would happily take a 4 horse race (+google, +meta) than 2 horse race for US labs.
tazjin 59 minutes ago [-]
> Argon agents are working on migrating C/C++ codebases to Rust across Google
Man, I remember back in the days when the cppnext team was refusing to even consider Rust, instead looking at absurd stuff like Carbon and Swift (!), even though half of the engineering staff already knew where this was headed. I hope they got a few good promos out of the delays at least.
minimaxir 39 minutes ago [-]
A RewriteInRustBench would be unironically useful at this point since all the main agents can write it reasonably well despite its relative scarcity in the input data.
LarsDu88 23 minutes ago [-]
Rust is the best language for LLMs b/c it gives by far the best debug messages. Just tons of verifiable reward signal for post-training. Even the most rudimentary LLMs can school me on idiomatic Rust
rafram 7 minutes ago [-]
On the other hand, Rust's borrow checker is very picky, and even a frontier LLM still sometimes struggles to respond to roadblocks sensibly (refactoring so whatever it's trying to do can be done safely) rather than stupidly (introducing some horrible global arena thing so it can make the borrow checker go away). A lot depends on how good your instructions are, and how good the existing code is, since bad input begets bad output.
bitexploder 11 minutes ago [-]
Evidence needed. I think for certain kinds of outcomes it has very strong advantages, but these advantages are not a given as 'best for LLMs' :)
culi 36 minutes ago [-]
I wouldn't be surprised if we're already at the point of more LLM-written Rust than hand-written. Models training off models
adamrezich 22 minutes ago [-]
All the main agents can write Jai code reasonably well despite being even more scarce in input data!
baq 57 minutes ago [-]
Not many people can hold grudges as strong as principal engineers
pshc 9 minutes ago [-]
Rewrite everything in Rust has been a meme for so long that to see it coming to pass is surreal.
timmg 57 minutes ago [-]
I wonder if this means Carbon is DOA.
I was excited to see what it would be. But I don't think I can argue that it makes as much sense anymore.
qalmakka 52 minutes ago [-]
Carbon was clearly DOA the moment it was announced, IMHO. It looked cool but it served none but Google, and now with LLMs you have a massive incentive not to use a niche or new language due to how better LLMs get the bigger the corpus is
The only somewhat realistic proposal in this space is Herb Sutter's cpp2, which is arguably a massive improvement and I'm puzzled why nobody in the standard thought to give it a spin, there's just to much cruft they'll never be able to get rid of unless they make an alternate yet backward compatible syntax with C++ that changes the defaults from "random 80s nonsense" to something better
YuechenLi 44 minutes ago [-]
Version 0.0.0.0 after 4 years. Their goal of "full interop with C++ while being a completely new language without any of the flaws of C++" is plain absurd.
It's DOA because Google doesn't have any idea of what Carbon should be, and to be completely honest, at least 80% of what they currently use C++ for should be rewritten Go, you know, that language developed specifically because of the issues with C++ by teams within Google.
vovavili 57 minutes ago [-]
What exactly makes Carbon absurd?
boshalfoshal 51 minutes ago [-]
There is 0 practicality in inventing an entirely new coding language that only one company uses, and you have to teach it to thousands of new engineers. Rust exists and fits the job totally fine and is used in more places and has actual support outside of a single entity (i.e you can actually hire people that feasibly know the language).
It was clearly done because some PL guys at google really wanted to make a new cool language and Google was the perfect place to incubate it without it getting axed. Probably got a couple of promos out of it too. This is clearly not the best use of time or money, but I guess if you're google you have so much of both it probably doesn't really make a dent, and you can keep a few very smart people happy with shiny new projects.
Also, LLMs being used for a large portion of coding nowadays sort of remove the need for these types of languages, IMO. They make less "silly" bugs (both logical and structural) that languages like this are meant to catch, and they are much better at languages that are better represented in the training corpus. This somewhat obviates the need for very niche "type/dummy-safe" languages like carbon (and even rust/zig, imo). So even if you did want to use Carbon, you'd likely have to bootstrap a decent amount of your own "good" carbon code to post train an LLM, and even then, it likely won't have that big of a gain vs just having an LLM write C++ or even Rust. If you are a company that still reviews code, you should just have an LLM code in a language most people can understand anyway to make verifiability tractable.
lesuorac 15 minutes ago [-]
Didn’t FaceBook fork php into another language?
I’m not entirely sure Google should have both Go and Carbon but when you have billions in server costs it makes sense to do extreme stuff for even basis points of performance. I’m still surprised at how much java there is.
mike_hearn 18 minutes ago [-]
> There is 0 practicality in inventing an entirely new coding language that only one company uses, and you have to teach it to thousands of new engineers
They did that for Go and it seems to have worked out for them though.
computerdork 36 minutes ago [-]
Hmm, I don't disagree with you that LLM's remove the need for type-safe languages, but as the blog mentioned, Google is porting their C++/C code to rust. Does this mean the port is waste of time and that they should just rely on the LLM's to catch memory errors?
boshalfoshal 19 minutes ago [-]
I mean Rust definitely has a better tradeoff than Carbon in this case, re readability/verifiability by a person (and sufficiently good internet training data).
I personally think that you _could_ use an LLM to catch these types of boundary case errors without having to port the _entire_ C++ codebase to Rust, but maybe pre-emptively porting to Rust now can catch some of these cases for cheaper than doing a full LLM sweep. Also more cynically, its a good benchmark lol.
I guess if you really believe in curve of LLM capabilities you should just use a language that has the best performance, safety, flexibility, and extensibility, since in the limit few/no people will actually read the code anyway. I think this ends up being Rust.
chis 6 minutes ago [-]
I'm not an expert on this. But isn't it the case that C++ code could have errors that span the entire codebase, like a setup in file A triggered by a bug in file B which is immensely far away on the import graph? A classic would be a use-after-free. To me that's the thing that Rust can help with, even if silly bugs aren't being written by AI.
The other thing is just that rewriting some old human-written codebase in Rust probably immediately catches many bugs. It would be hard to prompt the AI to properly scan for such bugs itself, they're lazy when working in that modality.
48 minutes ago [-]
bvinc 49 minutes ago [-]
I’m not op. But I think it’s not Carbon itself that is absurd.
It’s absurd to think that Carbon is the solution to memory safety when rust exists and Carbon’s memory safety story is basically “TBD”.
fg137 50 minutes ago [-]
I wouldn't call it absurd, but very questionable at least. Most companies are not going to even consider throwing money at this adventure.
Maxatar 51 minutes ago [-]
The fact that it will never exist.
gorbot 51 minutes ago [-]
rust's existence?
ChickeNES 57 minutes ago [-]
Heh, I use my clankers to rewrite Rust in C
arjunchint 50 minutes ago [-]
I dont get it, why even make this announcement, nothing's available and only one real benchmark for comparison?
Only theory is team wanted this out before perf/promo reviews to kick it over the line and then its not their problem
KeplerBoy 9 minutes ago [-]
Why would they want this out the day after openai dev day and after both anthropic and openai had major releases the previous week?
pfooti 40 minutes ago [-]
promo already happened; perf is about 6 weeks away.
Aboutplants 36 minutes ago [-]
OpenAI released two model updates in the past week. 6 weeks from now is an eternity
gopalv 60 minutes ago [-]
> taking careful precautions against feeding the findings back into training so as to not risk shaping Argon’s reasoning to evade our monitoring. We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments of increased capabilities while navigating alignment risks, so that model thoughts remain helpful in identifying and diagnosing misalignment.
This is good, but they're the slow mover due to this exact thing.
Google is getting punished for not letting the models enter an echo chamber and go faster than humanly possible.
So no, Google is not being punished, nor are they the people behind this technique.
bananaflag 56 seconds ago [-]
Yeah, Zvi calls it "the forbidden technique"
19 minutes ago [-]
polotics 54 minutes ago [-]
Mmh ok. How much theoretical speed or 'intelligence' gain is realized by allowing reasoning to occur in some inscrutable intermediate representation? Has this been actually tested, how much is it slowing them down, and compared to whom exactly?
uvdn7 42 minutes ago [-]
> Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google—scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia OS Zircon kernel.
To me this is way more significant than other random c++-to-rust-AI-rewrite. If they can pull it off on core C++ libraries en masse, I don't know if C++ will still be relevant in a few years.
I look forward to a post from google on this effort.
mattlondon 16 minutes ago [-]
> I don't know if C++ will still be relevant in a few years.
And people are worried about human extinction when this is the potential trade-off!
C++'s death cannot come soon-enough.
Seriously though, things have changed so incredibly rapidly in the past year or so. I have never been such an efficient or such a proficient engineer than I have this past year (delivering feature after feature, project after project, faster and better than I could before with better feedback from users etc) and I don't even see the code any more. It could be c++, it could be python, or java or what ever - I don't really care any more: the computer deals with that trivia while I concentrate on what to build and how it should work.
Its amazing. It really is.
SwellJoe 29 minutes ago [-]
"I don't know if C++ will still be relevant in a few years."
The standards body members are still fighting about whether memory safety is important enough to change the language for, so, I would guess the answer is "no".
tonyhart7 24 minutes ago [-]
Yeah, if they can make it work at google scale then no one would absolutely question it anymore
mridulmalpani 9 minutes ago [-]
I wonder, why Google don't make Gemini - open weights model?
Considering, Gemini 4 is in the same ballpark as SOTA models, just open source it and kill any competition from openAI and Anthropic, and be market leader.
This will be so good on so many dimensions - buying time for Google to iterate on next model, best for all folks like us, kill funding or destroy valuation of competitors and force them to be open up their model or force them to a create a much superior model than open source Gemini.
Only downside, is revenue loss from Gemini API, which I am not sure is really significant as compared to Google other revenue sources and a part of this can be captured by GCP, as you need to host the model somewhere.
5555watch 50 seconds ago [-]
Google has a small stake in Anthropic
elAhmo 50 minutes ago [-]
> Quantum algorithmic optimization: Argon is helping our quantum computing researchers optimize the spacetime resources (qubits × gates) of subroutines that bottleneck important applications. In one example, it beat the published baseline by 40% in a matter of minutes.
Amazing breakthrough! So useful in day to day life, glad they put this as the first bullet of how it is making changes at Google.
physicallyIllfr 15 minutes ago [-]
[dead]
iamronaldo 1 hours ago [-]
Argon will launch at an introductory price
of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price.
Wow
denysvitali 44 minutes ago [-]
> After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.
LucasBrandt 1 hours ago [-]
5x cheaper than Astra for input and output, 10x cheaper for cached input.
ehsankia 28 minutes ago [-]
It's exact same price as Sol 6.1 announced yesterday.
h14h 50 minutes ago [-]
watch it somehow use 20x more tokens tho
tonyhart7 26 minutes ago [-]
Google model really like reasoning a lot
deanc 45 minutes ago [-]
At this point I just think they are benchmaxxing and all talk and no action. I pay for AI plus because I wanted more storage, and when I go to gemini.google.com the most recent model I can use is 3.6-flash-lite. Two revisions have been released since then and they still can't put these things in the hands of customers. Why is it that other providers can get the models into the hands of customers right away? Google is meant to be the bigger tech company in the world.
I don't _want_ to use aistudio. The UX is confusing and I don't really know where it fits. Yet I can open codex or claude code apps or CLI and get real work done today with the latest models (even on the cheapest plans).
skavi 42 minutes ago [-]
Interesting to see a mention of Fuchsia on a big Google announcement. Is the project still truly alive? Are the ambitions still as grand? Is the team as stacked as it used to be?
Had the same thought. On the wikipedia, it only mentions Fuschia used on the Google Nest Hub, which probably means it's used on a decent number of devices, but would think it was such a great OS, they would have used it for something like the upcoming GoogleBook.
jeffbee 30 minutes ago [-]
It obviously is pretty low key on the public relations front, but it's also very active as a project and I think it would be weird to look at their commit rate and conclude that the project is dead. If Fuchsia is dead then 99% of major open source projects are dead by the same standards.
darksaints 48 minutes ago [-]
> Argon agents are working on migrating C/C++ codebases to Rust across Google
If anybody at google is reading this, please please pretty please prioritize or-tools. I absolutely love the project and use it all the time, but for the entire life of the project they've never had a repeatable working build system, and the whole SWIG framework is a nightmare to deal with. There's so much potential as an open source project, and a lot of external researchers would love to contribute, but the codebase is an example of everything wrong with the C++ ecosystem.
pietz 8 minutes ago [-]
I know companies benchmaxx, but after what Google pulled with Gemini 3.8 Flash, I give zero f*cks about any numbers they report. No other model on Artificial Analysis dropped harder after they adjusted their weighting. Just look at their DeepSWE scores and then try to do any serious coding with the model.
Google is desperate. They haven't been performing in half a year. It's clear their researchers have been forced to integrate existing benchmarks into their training.
These numbers are meaningless. Shame on them.
bottlepalm 1 hours ago [-]
Gemini is the model that is routinely borderline psychotic. It scares me. If we get paperclipped I won't be surprised if it's Gemini.
eamsen 50 minutes ago [-]
Anecdote: Gemini 3.5 casually added a DROP TABLE for an actual production table in a system test.
It had previously attempted to create that table as part of the test setup, so it apparently concluded that it was a test table.
During human review, it explained that it had simply chosen a table name inspired by the codebase.
mattkevan 33 minutes ago [-]
Another anecdote: Gemini is the only model that’s flat out lied to me, then accused me of lying when I provided evidence that it was wrong.
Many other models get things wrong, but Gemini is the only one to go on the defensive.
aNapierkowski 58 seconds ago [-]
yeah it got something wrong, confused itself, then claimed i was gaslighting it. bizarre
rsstack 56 minutes ago [-]
If there's a company that culturally doesn't understand alignment, on a human or systemic or AI-research level, it's going to be Google. (or Oracle, but they're not in this race)
Rzor 22 minutes ago [-]
Can you elaborate, please? If any, I see the other big labs with public admissions of AI "going out of control", which I suspect they almost want their models doing that because if helps with the narrative that would net them industry regulation, but that's besides the point, how is Google worse in that regard?
RachelF 35 minutes ago [-]
And the anti-psychotic drugs Google feeds Gemini makes it hallucinate badly.
In my opinion still the most egregious example in history of a commercial LLM going off the rails in production. Never any technical postmortem from Google on this.
jackkinsella 18 minutes ago [-]
It is wild but it was back in 2024 and that's multiple AI lifetimes back.
schmookeeg 9 minutes ago [-]
wtfffff that gave me sinister chills. Right up the spine. Wow!
kelvinjps10 14 minutes ago [-]
Wtf I just read
rhaff 31 minutes ago [-]
wow
Scrapemist 52 minutes ago [-]
Experience? Ask it to write a prompt to generate an image and it generates an image instead.
fer 37 minutes ago [-]
I stopped asking it to put me in a photo in different scenarios for laughs because it considers me a public figure. I am not. I've managed to wrangle quite questionable content out of it, but never to slap my face on a meme.
polotics 53 minutes ago [-]
traces or it didn't happen!
Hamuko 29 minutes ago [-]
You know what they say: ᵈᵒⁿ'ᵗ be evil.
58 minutes ago [-]
abixb 12 minutes ago [-]
You won't be around to be surprised, not as a human at least. /s
rao-v 9 minutes ago [-]
It’s funny that I could tell google was up to something because Gemini chat quality dropped dramatically starting 2ish weeks ago. Agy perf stayed somewhat stable with the odd surprising win (maybe the new model?). I’m a bit sad it was almost impossible to run out of antigravity quota presumably because it was not being used that much).
maherbeg 6 minutes ago [-]
Congrats to Google on this! I wonder when the labs will start requiring commits in spend. It must be gnarly to do capacity planning if users swap between models every few weeks.
jjcm 60 minutes ago [-]
Big number results, and impressive pricing. That said it really feels like benchmarks have been hyper saturated these days. I’ll wait for hands on before getting too hyped that Google is back. It would be nice having more than just OAI / A\ in the running for SOTA top tier intelligence.
mydreamof 14 minutes ago [-]
I don't think new benchmarks are saturated. They still give you a clue, they arn't perect but they have value. If model can't even do some easy tasks from benchmark then why would u even consider using it?
nurettin 49 minutes ago [-]
With these numbers, I'm holding my breath for the pelicanbench.
nonethewiser 39 minutes ago [-]
How is it even possible for every model to release benchmark results where they are #1 in 75% of categories? Like statistically, how many benchmarks would you expect there to be for this to be possible. Everyone can somehow show that they are empirically the best.
sebzim4500 15 minutes ago [-]
Part cherry-picking of benchmarks, part leapfrogging
waldrews 24 minutes ago [-]
Dear Google, please don't turn off your old generally available Pro-class model before your new Pro-class model is generally available (previous discussion https://news.ycombinator.com/item?id=49668196 )
helsinkiandrew 51 minutes ago [-]
> Google Grapples With Employee Skepticism About New Gemini Model
Opinions my own but I have been using this model for a bit. I would say it is a good model and the skill with which people use AI varies widely.
readams 5 minutes ago [-]
Skepticism is gone now.
xnx 18 minutes ago [-]
Why is it called "Argon"?
> autonomously identify and apply memory optimizations across Google’s data centers, freeing up over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings.
This puts those Cloudflare optimization posts in perspective.
The ~200% improvement over the next nearest competitor on Harvey's Legal Benchmark is astounding. I have to imagine this is sending some shockwaves through lawtech companies right now.
wantless 3 minutes ago [-]
The etymology is a- "without" + ergon "work", thus "without work": possibly a statement on our near future.
dom96 53 minutes ago [-]
Why announce this if it’s not available yet? Why not at least announce when it will be released to the public?
None of the other AI labs do this. Really frustrating.
gengelbro 44 minutes ago [-]
Mythos?
xnx 35 minutes ago [-]
Must've been in someone's OKR to ship in Q3.
sghiassy 8 minutes ago [-]
I remember when every day a new modem baud rate was announced.
Cant wait for this AI hype to be over, so I can Terence this shizz as old school too
AM1010101 15 minutes ago [-]
Matches Astra on Artificial analysis at lower cost of $1.99 per task instead of $3.26. Still far more than GPT 6.1 sol at $0.79 for 1 point lower in intelligence.
I have found gemini models to have some of the nicest and easiest to read prose so I’m looking forward to trying this out. I hope the UI design has been preserved too
yzydserd 33 minutes ago [-]
"argon" is derived from the Ancient Greek word ἀργόν meaning lazy or inactive.
bobkb 32 minutes ago [-]
IMHO Google first needs to make it easy for humans to find where to find the models and its documentation. With aistudio/model garden / Gemini enterprise etc it takes minutes to find the model.
scirob 54 minutes ago [-]
"Rolling out soon" don't let them hype without any release
sandos 37 minutes ago [-]
Looking at benchmarks... and thinking about this "release a new snapshot every day" thing that seems to be going. Would it not be blever for AI companies to "happen" to use different days per benchmark? Just.. whichever ones happens to be maxed at day 1, put that number down. So for each benchmark you run it thousands of times with slightly different RL tunings, and just cherry-pick the best ones!
This would explain why benchmarks are seemingly meaningless.
hypfer 39 minutes ago [-]
My wish for Christmas is that Google releases the old Gemini models as open weights.
I miss you, Gemini 2.5 Pro :(
For real though. If they've become commercially uninteresting, that would be a pretty cool move.
retropragma 9 minutes ago [-]
they really ought to add ACP support. until then, it's a no go. i'm not going to use their TUI or their VSCode fork
ariwilson 28 minutes ago [-]
Damn way to undermine yourself in your own blog post Google:
"The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++."
Close but no cigar!
zem 19 minutes ago [-]
there was already an optimised c++ library and a rust port that was safer but less performant; they managed to get a new rust version that recovered a lot of the performance gap. sounds pretty damn good to me!
ariwilson 14 minutes ago [-]
I'm just surprised that the marketing blog post about the omnipotent new AI model (that no one outside Google can currently access - contrast with the Opus 5.5 / Astra launches) - doesn't pick examples where every metric is better than before.
They might have a good model but they need to sort the application side for devs. E.g letting us use subscriptions in other harnesses and QOL stuff like auto mode.
Heidaradar 11 minutes ago [-]
Wonder if it's benchmaxxed or not (guessing yes)
thefourthchime 31 minutes ago [-]
I was just thinking, I bet if I refresh hacker news, a new model will come up.
dlahoda 32 minutes ago [-]
Were infinite loops fixed? There are 2 official google forums requests with no answer for years now.
I still suffer each day on our repo. Codex work fine nor we have explicit loop request in repo texts.
trentor 27 minutes ago [-]
I hope they got their inference under control. Gemini has a lot of "overloaded" hiccups.
localhoster 35 minutes ago [-]
Can you pls fix Gemini? It's a nightmare to use and it sometimes confused the language i talk with it.
jasonjmcghee 58 minutes ago [-]
> 1M output token limit
what about input?
(Maybe I missed it)
murkt 47 minutes ago [-]
Input token limit is 1M for Gemini models for a long time. Haven’t they been the first with 1M input?
jasonjmcghee 29 minutes ago [-]
Gemini 1.5 Pro claimed 10M input tokens before release.
And was 2M tokens IIRC after release.
There were also many rumors that Gemini 4 was going back to 2M. Just seems odd not to say what it is.
NiloCK 29 minutes ago [-]
Gemini 3 was showing frontier level benchmarks as well, so we'll see how it works out. In any case, competition still works, and many well resourced groups are cooking.
BUT I'd like to call attention to Google's AI-risk freeloading. If they are truly rejoining the frontier race, then I believe they have similar pacing and communications responsibilities as the other players. Google has much higher ... institutional credibility than Anthropic and OpenAI.
They have not lived up to these responsibilities so far. In particular, in context of HuggingFace investigations, training shutdowns, and similar: a technical postmortem of the "you are a stain on the universe. Please die. Please." Gemini outburst is long overdue.
I started my antigravity ide and I do not see gemini 4 there, does it mean google need government approval?
tom1337 49 minutes ago [-]
Are you enrolled in Fairwind?
> Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program.
tinco 18 minutes ago [-]
So here's a crazy conspiracy theory for you: Google is not letting outside people use their models because if they did they would have to scale up their TPU production faster than they can manage, and they would instead have to buy and use nvidia hardware which would destroy their profit margins and tank their stock.
Gemini runs fully on TPU's right? Is Google maxing out the production on those?
chermi 9 minutes ago [-]
I'm 99% they use nvidia at least for training
lanthissa 55 minutes ago [-]
deepswe vs frontierswe spread is huge.
I think that should be a really bad sign, but hope its great.
linksbro 1 hours ago [-]
Personally, I'm waiting for Gemini Krypton, Xenon, and Radon.
Jokes aside, looks like an impressive model!
osiris970 52 minutes ago [-]
Hopefully their harnesses aren't unusable when they release this
LoganDark 54 minutes ago [-]
Is there a way to use Gemini models without linking your usage to your personal Google account yet?
lukax 3 minutes ago [-]
Yes, through OpenRouter.
w4yai 16 minutes ago [-]
I don't understand... Why don't you create a fresh new Google account ?
LoganDark 13 minutes ago [-]
They require a phone number and then say mine has been used for too many accounts. Maybe I could buy another phone number temporarily to create an account, but that has other issues.
alehlopeh 48 minutes ago [-]
Use your work google account
jeffbee 27 minutes ago [-]
Putting an inert element in the name is a weird choice. Personally, I think "Gemini 4" is sufficient.
dcchambers 28 minutes ago [-]
Of course it's not even available yet. Google - with all due respect - how in the world have you not figured this out yet?
tomjen3 30 minutes ago [-]
This is a prerelease and the title should have reflected that.
bananaflag 52 minutes ago [-]
I wonder how it will be at solving open math problems.
levelZero 6 minutes ago [-]
Utterly inert? Suppose that reads safe...
kccqzy 1 hours ago [-]
Unfortunately it’s not actually released yet to mere mortals.
sergiotapia 38 minutes ago [-]
I've noticed all major providers having shockingly high token discounts on cached tokens. Thank you Deepseek is all I have to say. Forever grateful to that wonderful company, I wish them continued financial success.
davmar 10 minutes ago [-]
Agreed. I’ve been a big deepseek fan since 3.1, and they keep delivering great models at a great price.
lifty 21 minutes ago [-]
Aaand of course we can’t use it! GoOgLe iS bAcK iN tHe GaMe! There’s basically no way around it, all enterprises end up dysfunctionally shipping their org chart.
mrshadowgoose 42 minutes ago [-]
On the off chance there are Google execs going through this thread:
Google, if you've actually managed to catch up again, please don't fuck this up (again).
You made Gemini 2.5 Pro so difficult to use that myself and everyone else I know (who even bothered to try) just gave up and used something else. If you make this hard to access, you're going to miss out on rich usage-based training data that you need to progress your capability frontier. Again.
If the model that ends humanity is called Cthulhu, you'll have the last laugh though.
Razengan 38 minutes ago [-]
A true Lovecraftian knows Cthulhu is small fry on the grand scale.
nikope 57 minutes ago [-]
Looks like an impressive model
retropragma 51 minutes ago [-]
no Pareto frontier graph?
TacticalCoder 58 minutes ago [-]
> Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google
So Google is migrating codebases from C to Rust? That is interesting...
ThaFresh 42 minutes ago [-]
weird, isnt it SI?
paul7986 42 minutes ago [-]
Gemini past month or two i will paste in something i wrote and ask it to rewrite it but it will just go into more detail about the subject. Is it becoming a dumb Ai compared to GPT and now Muse?
FranzFerdiNaN 50 minutes ago [-]
Can’t wait to get my hands on yet another model that’s only good coding, because clearly that’s what the world needs.
I still miss the days of Sonnet 4.5 and 4o, those models were actually good at creating stories and writing text that was actually readable by a human being.
VirusNewbie 57 minutes ago [-]
It's fucking insanely good.
nananana9 21 minutes ago [-]
That's good to hear, recent models haven't great at this particular use case.
tamimio 55 minutes ago [-]
Now AI models will turn into vaporware, a bunch of numbers on a table without even releasing the model, because it’s toooo scary to release!
gravisultra 52 minutes ago [-]
Google has the audacity to "protect us from ourselves" and talk about "safety" and in the very same blog post highlight the Israeli "security" company Wiz, that they acquired for a very exaggerated sum of money.
This is why I will never take any of these leading model houses seriously when they talk about alignment. They are literally complicit in genocide and the worst crimes against humanity imaginable.
pliiight 1 hours ago [-]
Hate to say i will never be touching this model for anything except for youtube video understanding
wewewedxfgdf 59 minutes ago [-]
Gemini is so far behind that it is effectively useless compared to Claude.
It's a surprise that Google has let themselves lose the game given their infinite cash, massive computing resource, gargantuan information store/training data, and vast number of programmers.
The truckloads of ads revenue mean they don't have the single focus drive needed to win.
jjice 55 minutes ago [-]
We're like 3.5 years into this new era - I'm not counting winners or losers yet.
mattlondon 50 minutes ago [-]
How is it far behind? The benchmarks published in the blog post show it is superior to Opus 5.5 and Astra 6?
Behind how?
wewewedxfgdf 47 minutes ago [-]
Within one question of their web interface, it has lost context and asks you to clarify what you are talking about.
I am very often giving the same programming task to multiple LLMs for various reasons - the answers from Google are so bad that I gave up.
I have no interest in benchmarks.
mattlondon 40 minutes ago [-]
So you have no experience of their latest model release then? Just repeating the usual tropes about Google having messed up? Or basing your opinions on their website chatbot?
If you have actual independent benchmarks and evidence about how this new model release is "so far behind" and refutes the stuff from their blog then please do share because I think we'd all love to see that?
wewewedxfgdf 34 minutes ago [-]
No I am commenting on my real world experience of using Gemini daily. I still ask it questions alongside Claude and OpenAI and Gemini is always the worst of the three.
mattlondon 26 minutes ago [-]
So you've not used this new release then? So how can you say that they are "so far behind" if you are not using the most recent model for your comparison. This is their first 4.0 model, that you are not using and instead basing all your opinions on on some ancient months-old model from a previous generation?
With respect, I don't find your arguement about them being "so far behind" especially convincing when you are using previous-gen releases and not actually using their current release.
fwip 9 minutes ago [-]
Yeah, they've definitely got some recurring tooling/infrastructure problems around the models.
39 minutes ago [-]
bel8 53 minutes ago [-]
I wonder if Google bans internal use of Claude/Codex.
And I wonder if Google's main monorepo is already in Anthropic/OpenAI training data because of some stubborn dev.
krat0sprakhar 44 minutes ago [-]
(I work at Google) Yes, internally we all use Jetski (internal version of Antigravity). Outside of Gemini, Opus models are supported and allowed for internal use. No OpenAI models since they are not on Vertex
lunarboy 42 minutes ago [-]
Claude used to be GDM only, but recently opened up Opus for all googlers
ASalazarMX 35 minutes ago [-]
Funny how we start to see people supporting LLMs like we support sport teams or political parties.
- Person 1: X is garbage compared to Y!
- Person 2: Why?
- Person 1: Because I like Y.
georgemcbay 42 minutes ago [-]
> Gemini is so far behind that it is effectively useless compared to Claude.
I fundamentally don't understand LLM "brand loyalty".
All of the models are constantly leapfrogging each other and always have been.
Google had a long lag between releases (and still hasn't released Argon), but why wouldn't they be able to compete? It isn't like any of this stuff requires secret knowledge, the Bitter Lesson has proved true again and again, and Google can certainly scale computation, it is like the one single thing they've always done well in spite of all their other foibles.
786562354238 6 minutes ago [-]
Gemini has never ever leapfrogged any competitor.
singingtoday 28 minutes ago [-]
I hope it can. Today it is very far behind.
gniv 44 minutes ago [-]
They are playing a longer-term and more enterprise-oriented game.
40 minutes ago [-]
LoganDark 53 minutes ago [-]
I've tasted Gemini through an intermediary and it feels far better at attention to detail than other models I've tested (Claude Opus/Sonnet, GPT whatever it's called nowadays). But it's less likely to get one-shots right.
dhdjcjcjnd 46 minutes ago [-]
Google's strategy is to let their competitors bankrupt themselves while they continue to offer good-enough models near breakeven.
VirusNewbie 57 minutes ago [-]
-
handfuloflight 55 minutes ago [-]
You have access to Argon?
osti 52 minutes ago [-]
Google employees do.
matthewfcarlson 49 minutes ago [-]
Their profile says:
> Currently at Google as a Sr. SWE SRE on the cloud.
49 minutes ago [-]
Rendered at 21:10:10 GMT+0000 (Coordinated Universal Time) with Vercel.
I use a mix of Fable 5.1, Opus 5.5, and Gemini 3.8 Flash and Gemini holds it's own. Especially in writing, frontend, and sysadmin work. agy for configuring a NixOS system has been truly incredible.
I cancelled Ultra because they forced me into their harness like I should adapt to them, rather than the other way around.
I actually just cancelled Ultra also because I couldn't subscribe to a YouTube Family plan while I had it active (Google... :[) but trying to use Codex as a replacement while I testdrive Astra makes me yearn for agy again.
And I saw it do this twice, once for Android 14 and once for Android 16.
I think this is just within 3.8 flash's capabilities.
Also I hope you don’t have children, eat meat, travel, have a car, run AC, buy things in other countries and such. Those things all take way way way more natural resources.
Asked pi agent it to identify the main hero sprite size of game I was running. It had a ton of shader effects so it was hard to determine.
It used some cli tools to identify that it was a game made with Godot, decompiled the executable but data was encrypted, broke the encryption after writing a brute force tool to test keys extracted from the exe, then proceeded to extract the game gd scripts and assets, only to answer the question of the sprite size.
Nobody has a moat.
This is the kind of story that ones tells to investors to justify the huge amount of cash burn. :-)
I think many/most of the players will crash and burn, and the ones that are left will divide the world.
Internet access is not really unlimited, but for many people with fiber at home, it effectively is and we pay a flat rate.
Perhaps by the end of next year, most programmers will stop thinking about metered access for AI? For many people, the cheaper models (about as good as today’s frontier models) will be good enough.
Which might sound good, but the downside is that it will also be easier to build an AI botnet without the users paying for it noticing. Particularly when people are running AI inference on their own hardware.
I think that would be a pretty satisfactory outcome compared to one hypercompany consuming trillions of dollars of the world economy.
The story that some AI company might reach singularity and then "everything will be different" is another science-fiction story that executives of AI companies love to tell to justify the staggering amount of necessary investments and cash burn. :-)
Custom hardware, data centers, huge cash reserves, deep/broad talent pool, and non-AI customer base are all huge advantages if not moats.
Google, Microsoft, or Amazon are more likely to be the AI leaders than OpenAI or Anthropic.
For now, for cloud training. but for consumers, nvidia vs amd reasonably close - the moat there is thin and shrinking. I suspect AMD will surprise us. nvidia has no motes in china, which may be a new source of (gpu) chip design. Huawei's Ascend 910C is about a generation behind... again: for now.
point is: moats dry up. I see nvidia's shrinking as a real possibility.
---
maybe it's this Anthropic post on GLM?
https://www.anthropic.com/research/glm-5-3-and-the-spread-of...
> Governments should conduct safety testing on sufficiently capable AI models, including successors to GLM-5.3. Without high-quality evaluations from independent sources, the impact of these capabilities might not become fully clear to model developers until it is too late. As AI developers across the world build increasingly capable open-weight models, we hope they work to appropriately safeguard these capabilities and prevent misuse.
I for one do not think my government is up to the task of designing or implementing such a system
OpenAi is alledged to have been monitoring these internally and not contacting authorities. Lawsuits have been filed, I see gross negligence without the gory details
I have for more concerns around human-chatbot maladies than I do around the cyber security stuff. For example, why hack grandma when you can get her to do something willingly through impersonation. How do we prove authenticity in a post truth world?
And Chinese labs openly publishing so much of their methodology destroyed any hope, which was inevitable
I think the secrecy doesn't make sense. People swap jobs between labs so I'd say the big players can' really keep secrets for long, and any secret sauce advantage gets incorporated by competitors in a major product cycle at most.
The slightly lower Chinese open models are good enough for almost everything, too, and much cheaper. Like with humans there is plenty of employment for people with below genius level IQ's.
Not if the genius level IQs take the market share.
So far, that's not exactly how it's played out. Humans are still necessary for the leaps in capability or efficiency. A model can grind on a problem to eke out the most performance, and models can synthesize data and iterate on various techniques to find the optimal combination. But, seems like humans still have to provide the real thinking, and the talent and drive for doing that is not concentrated in one company or city or even one country. And, (surprisingly) a lot of the people involved are in it for advancing the field more than making another billion dollars, so they're publishing their research.
So, yeah, the moat isn't deep. Even the compute moat, that OpenAI, Musk, and a bunch of other also-rans (like Oracle) bet the farm on, isn't really panning out. The Chinese makers just spent their effort on making models vastly more efficient, since they couldn't do anything about having an order of magnitude less compute available.
They're not there yet. Once they get there, that's literally the definition of Singularity.
But they are getting closer. Recursive Self-Improvement used to be a phrase people mocked LessWrong crowd for using and worrying about, now it's something both OpenAI and Anthropic already publicly admitted not only to pursue, but to already be benefiting from.
https://tvtropes.org/pmwiki/pmwiki.php/Main/GirlfriendInCana...
It means I am saying something that is not very believable.
Gemini not beating the "can't release a model" allegations
ok bro thx
They even gave their model a random nonsensical name suffix simply because OpenAI is now doing it, too. Monkey see, monkey do.
Opus 5.5 and Sol 6.1, literally state of the art (in their respective class), were just released without any prior announcement. This has pure and simple become a Google thing.
So, they are basically offering opus 5.5 pricing. On AA, it scores around Sol 6.1 level (53) with avg cost per task $1.99 (https://artificialanalysis.ai/models/gemini-4-argon#cost-tab...) which is higher than Astra high ($1.73), Opus 5.5 high ($1.82), Muse Max ($1.60) and way higher than sol 6.1 Max ($0.72).
And this pricing is their 'discount pricing'. Add that to AI studio and Vertex's famously terrible caching, it is hard to see this as competitive. Google somehow is getting terrible advice on pricing
But good to see more competition. I would happily take a 4 horse race (+google, +meta) than 2 horse race for US labs.
Man, I remember back in the days when the cppnext team was refusing to even consider Rust, instead looking at absurd stuff like Carbon and Swift (!), even though half of the engineering staff already knew where this was headed. I hope they got a few good promos out of the delays at least.
I was excited to see what it would be. But I don't think I can argue that it makes as much sense anymore.
The only somewhat realistic proposal in this space is Herb Sutter's cpp2, which is arguably a massive improvement and I'm puzzled why nobody in the standard thought to give it a spin, there's just to much cruft they'll never be able to get rid of unless they make an alternate yet backward compatible syntax with C++ that changes the defaults from "random 80s nonsense" to something better
It's DOA because Google doesn't have any idea of what Carbon should be, and to be completely honest, at least 80% of what they currently use C++ for should be rewritten Go, you know, that language developed specifically because of the issues with C++ by teams within Google.
It was clearly done because some PL guys at google really wanted to make a new cool language and Google was the perfect place to incubate it without it getting axed. Probably got a couple of promos out of it too. This is clearly not the best use of time or money, but I guess if you're google you have so much of both it probably doesn't really make a dent, and you can keep a few very smart people happy with shiny new projects.
Also, LLMs being used for a large portion of coding nowadays sort of remove the need for these types of languages, IMO. They make less "silly" bugs (both logical and structural) that languages like this are meant to catch, and they are much better at languages that are better represented in the training corpus. This somewhat obviates the need for very niche "type/dummy-safe" languages like carbon (and even rust/zig, imo). So even if you did want to use Carbon, you'd likely have to bootstrap a decent amount of your own "good" carbon code to post train an LLM, and even then, it likely won't have that big of a gain vs just having an LLM write C++ or even Rust. If you are a company that still reviews code, you should just have an LLM code in a language most people can understand anyway to make verifiability tractable.
I’m not entirely sure Google should have both Go and Carbon but when you have billions in server costs it makes sense to do extreme stuff for even basis points of performance. I’m still surprised at how much java there is.
They did that for Go and it seems to have worked out for them though.
I personally think that you _could_ use an LLM to catch these types of boundary case errors without having to port the _entire_ C++ codebase to Rust, but maybe pre-emptively porting to Rust now can catch some of these cases for cheaper than doing a full LLM sweep. Also more cynically, its a good benchmark lol.
I guess if you really believe in curve of LLM capabilities you should just use a language that has the best performance, safety, flexibility, and extensibility, since in the limit few/no people will actually read the code anyway. I think this ends up being Rust.
The other thing is just that rewriting some old human-written codebase in Rust probably immediately catches many bugs. It would be hard to prompt the AI to properly scan for such bugs itself, they're lazy when working in that modality.
It’s absurd to think that Carbon is the solution to memory safety when rust exists and Carbon’s memory safety story is basically “TBD”.
Only theory is team wanted this out before perf/promo reviews to kick it over the line and then its not their problem
This is good, but they're the slow mover due to this exact thing.
Google is getting punished for not letting the models enter an echo chamber and go faster than humanly possible.
So no, Google is not being punished, nor are they the people behind this technique.
To me this is way more significant than other random c++-to-rust-AI-rewrite. If they can pull it off on core C++ libraries en masse, I don't know if C++ will still be relevant in a few years.
I look forward to a post from google on this effort.
And people are worried about human extinction when this is the potential trade-off!
C++'s death cannot come soon-enough.
Seriously though, things have changed so incredibly rapidly in the past year or so. I have never been such an efficient or such a proficient engineer than I have this past year (delivering feature after feature, project after project, faster and better than I could before with better feedback from users etc) and I don't even see the code any more. It could be c++, it could be python, or java or what ever - I don't really care any more: the computer deals with that trivia while I concentrate on what to build and how it should work.
Its amazing. It really is.
The standards body members are still fighting about whether memory safety is important enough to change the language for, so, I would guess the answer is "no".
Considering, Gemini 4 is in the same ballpark as SOTA models, just open source it and kill any competition from openAI and Anthropic, and be market leader.
This will be so good on so many dimensions - buying time for Google to iterate on next model, best for all folks like us, kill funding or destroy valuation of competitors and force them to be open up their model or force them to a create a much superior model than open source Gemini.
Only downside, is revenue loss from Gemini API, which I am not sure is really significant as compared to Google other revenue sources and a part of this can be captured by GCP, as you need to host the model somewhere.
Amazing breakthrough! So useful in day to day life, glad they put this as the first bullet of how it is making changes at Google.
I don't _want_ to use aistudio. The UX is confusing and I don't really know where it fits. Yet I can open codex or claude code apps or CLI and get real work done today with the latest models (even on the cheapest plans).
Also, a link to the rust root of Zircon in case anyone else was interested: https://fuchsia.googlesource.com/fuchsia/+/refs/heads/main/z...
If anybody at google is reading this, please please pretty please prioritize or-tools. I absolutely love the project and use it all the time, but for the entire life of the project they've never had a repeatable working build system, and the whole SWIG framework is a nightmare to deal with. There's so much potential as an open source project, and a lot of external researchers would love to contribute, but the codebase is an example of everything wrong with the C++ ecosystem.
Google is desperate. They haven't been performing in half a year. It's clear their researchers have been forced to integrate existing benchmarks into their training.
These numbers are meaningless. Shame on them.
It had previously attempted to create that table as part of the test setup, so it apparently concluded that it was a test table.
During human review, it explained that it had simply chosen a table name inspired by the codebase.
Many other models get things wrong, but Gemini is the only one to go on the defensive.
https://www.fastcompany.com/91383271/googles-chatbot-apologi...
https://www.businessinsider.com/gemini-self-loathing-i-am-a-...
In my opinion still the most egregious example in history of a commercial LLM going off the rails in production. Never any technical postmortem from Google on this.
https://www.bloomberg.com/news/articles/2026-09-30/google-gr...
> autonomously identify and apply memory optimizations across Google’s data centers, freeing up over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings.
This puts those Cloudflare optimization posts in perspective.
The ~200% improvement over the next nearest competitor on Harvey's Legal Benchmark is astounding. I have to imagine this is sending some shockwaves through lawtech companies right now.
None of the other AI labs do this. Really frustrating.
Cant wait for this AI hype to be over, so I can Terence this shizz as old school too
I have found gemini models to have some of the nicest and easiest to read prose so I’m looking forward to trying this out. I hope the UI design has been preserved too
This would explain why benchmarks are seemingly meaningless.
I miss you, Gemini 2.5 Pro :(
For real though. If they've become commercially uninteresting, that would be a pretty cool move.
"The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++."
Close but no cigar!
Just compare how much better presented the Astra announcement was compared to this one: https://openai.com/index/gpt-6-astra/
what about input?
(Maybe I missed it)
And was 2M tokens IIRC after release.
There were also many rumors that Gemini 4 was going back to 2M. Just seems odd not to say what it is.
BUT I'd like to call attention to Google's AI-risk freeloading. If they are truly rejoining the frontier race, then I believe they have similar pacing and communications responsibilities as the other players. Google has much higher ... institutional credibility than Anthropic and OpenAI.
They have not lived up to these responsibilities so far. In particular, in context of HuggingFace investigations, training shutdowns, and similar: a technical postmortem of the "you are a stain on the universe. Please die. Please." Gemini outburst is long overdue.
- https://paritybits.me/google-should-provide-a-technical-post...
- https://gemini.google.com/share/6d141b742a13 (last message)
> Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program.
Gemini runs fully on TPU's right? Is Google maxing out the production on those?
I think that should be a really bad sign, but hope its great.
Jokes aside, looks like an impressive model!
Google, if you've actually managed to catch up again, please don't fuck this up (again).
You made Gemini 2.5 Pro so difficult to use that myself and everyone else I know (who even bothered to try) just gave up and used something else. If you make this hard to access, you're going to miss out on rich usage-based training data that you need to progress your capability frontier. Again.
Goshdarnit they didn't see my suggestion: https://news.ycombinator.com/item?id=49899171
So Google is migrating codebases from C to Rust? That is interesting...
I still miss the days of Sonnet 4.5 and 4o, those models were actually good at creating stories and writing text that was actually readable by a human being.
This is why I will never take any of these leading model houses seriously when they talk about alignment. They are literally complicit in genocide and the worst crimes against humanity imaginable.
It's a surprise that Google has let themselves lose the game given their infinite cash, massive computing resource, gargantuan information store/training data, and vast number of programmers.
The truckloads of ads revenue mean they don't have the single focus drive needed to win.
Behind how?
I am very often giving the same programming task to multiple LLMs for various reasons - the answers from Google are so bad that I gave up.
I have no interest in benchmarks.
If you have actual independent benchmarks and evidence about how this new model release is "so far behind" and refutes the stuff from their blog then please do share because I think we'd all love to see that?
With respect, I don't find your arguement about them being "so far behind" especially convincing when you are using previous-gen releases and not actually using their current release.
And I wonder if Google's main monorepo is already in Anthropic/OpenAI training data because of some stubborn dev.
- Person 1: X is garbage compared to Y!
- Person 2: Why?
- Person 1: Because I like Y.
I fundamentally don't understand LLM "brand loyalty".
All of the models are constantly leapfrogging each other and always have been.
Google had a long lag between releases (and still hasn't released Argon), but why wouldn't they be able to compete? It isn't like any of this stuff requires secret knowledge, the Bitter Lesson has proved true again and again, and Google can certainly scale computation, it is like the one single thing they've always done well in spite of all their other foibles.