The BSDs have had a version of this for many, many years. If you hit ^T on the terminal, you deliver a SIGINFO to the application you're waiting for. By default, you get the program name, what its blocked on, and info about real/user/sys time and memory use.
It would be much better if we adopted SIGINFO + env variable where to respond.
This way, anyone who needs AI agent status can just send SIGINFO signal and wait for the response.
paaloeye 23 minutes ago [-]
Darwin also supports it
weinzierl 1 hours ago [-]
I've been using a poor mans version of this for decades.
My iTerm2 is configured to show activity, new-output and visual bell in the tab. On Linux I have an approximation for WezTerm.
I have a bell command that I can use in a pipeline or sequence to produce the bell on events I'm interested in. Trivial case is when a program finishes.
I have a fancy alias that can be used in a shell command sequence and which produces different sounds depending on the exit status of the preceding command in addition to sending the terminal bell.
zephraph 48 minutes ago [-]
I absolutely love this and hope it takes off. I've worked on several products with embedded terminals and reliably understanding when they were waiting for human input was such a pain.
flopsamjetsam 1 days ago [-]
Brings back to mind the old IBM 3270 terminal status line (everything old is new again :). Actually think this is quite a good idea.
paaloeye 7 hours ago [-]
> Actually think this is quite a good idea.
That's what I thought in the beginning. Once I read the spec to the end, it's clear to me that it's overly specific to one use-case — TUI AI agents.
E.g. `kind := permission | question | auth` doesn't make much sense anywhere outside of that use case. The list goes on...
I also found comms around it super muddy, AI agent use case is mentioned but deceivingly watered down with other use-cases.
Overall, that spec is no where close to Kitty/Kovid's level. It's super sad, since mitchellh's work is usually super high quality.
I hope the community will push back and the spec will get better.
cpuguy83 2 hours ago [-]
A cli asking for a password is a pretty normal thing, I think?
hinkley 59 minutes ago [-]
As a frequent internal tool writer, I’d say it’s a fact of life but an undesirable one. I’m always trying to avoid pausing for credentials in the middle of a run, either by doing them right at the beginning or by using some sort of durable credentials system like ssh keys.
There’s always two people at every company who refuse to set up auth automation though, so you end up handling it. They are very stubborn. Even when it trips them up in front of observers while executing an urgent task, I’ve only occasionally convinced them to agree to set up trust instead of typing a password every single time.
Which is to say, I have to support password prompts even though they drive me up the wall.
paaloeye 58 minutes ago [-]
Yes, but a CLI can ask for literally anything. We'll end up adding things there and with a zoo of hard-to-support implementation of that _protocol_.
CLI can also _blank_, do we want to support it as well?
kevin_thibedeau 1 hours ago [-]
It's a product of the scope limited interface granted to agents. They get a terminal stream so everything goes into the stream. Piling on more in-band signaling is just going to become a security nightmare. It would be better to have a safe way to query process state that can be locked down as needed.
danudey 1 hours ago [-]
As someone who works on a lot of CLI tools at work I actually like the idea of implementing this in some of our long-running tooling, completely independent of the AI agent use case.
The original idea I had being that CI environments, or anywhere that shows terminal output, could use a streaming ANSI parsing library to detect when a program sent a 'change window title' OSC event and then start a new collapsable section in its output. This would let programs update the actual terminal window title with its current state when running locally and update the CI interface when running in CI.
Adding in the ability to specify the program's current status and (optionally) progress could be extremely useful in CI environments as well. Imagine, for example, a long-running analysis task which uses OSC 7501 to say that it's currently running and is 75% of the way done. CI could expose that in the UI to provide useful information to the end-user without having to print a progress bar or multiple progress lines to the fake terminal it's being run in.
paaloeye 33 minutes ago [-]
long-running analysis task can use osc9;4 for progress reporting
My first brush with CI, I ended up putting an option to play a sound at the end of local builds because I’d already noticed evidence of Hofstadter’s Law applying to build automation.
The thing is when you expect a task to take five minutes, you don’t watch it, you find something else you expect to take five minutes and do that instead. When that ends up taking ten minutes, or when you remember what you were doing before you started, you finally come back around ten minutes later to find that either the task completed four minutes ago, or it failed after ten seconds and you’ve wasted ten now.
The audio was the best out of band notification I had at my disposal 20 years ago.
My first thought when reading this was actually terminal multiplexing however, like screen or tmux. But I’m also always doing the terminal dance because I work on 4 FOSS projects and I keep 1+ terminal open per project so I can jump in and do bug fixes or pull PRs I’ve landed.
Starting an emacs window in the fg:
% emacs ^T load: 0.32 cmd: emacs-31.1 64938 [select] 2.89r 0.75u 0.06s 6% 122404k
This way, anyone who needs AI agent status can just send SIGINFO signal and wait for the response.
My iTerm2 is configured to show activity, new-output and visual bell in the tab. On Linux I have an approximation for WezTerm.
I have a bell command that I can use in a pipeline or sequence to produce the bell on events I'm interested in. Trivial case is when a program finishes.
I have a fancy alias that can be used in a shell command sequence and which produces different sounds depending on the exit status of the preceding command in addition to sending the terminal bell.
That's what I thought in the beginning. Once I read the spec to the end, it's clear to me that it's overly specific to one use-case — TUI AI agents.
E.g. `kind := permission | question | auth` doesn't make much sense anywhere outside of that use case. The list goes on...
I also found comms around it super muddy, AI agent use case is mentioned but deceivingly watered down with other use-cases.
Overall, that spec is no where close to Kitty/Kovid's level. It's super sad, since mitchellh's work is usually super high quality.
I hope the community will push back and the spec will get better.
There’s always two people at every company who refuse to set up auth automation though, so you end up handling it. They are very stubborn. Even when it trips them up in front of observers while executing an urgent task, I’ve only occasionally convinced them to agree to set up trust instead of typing a password every single time.
Which is to say, I have to support password prompts even though they drive me up the wall.
CLI can also _blank_, do we want to support it as well?
I actually created a separate golang library with something like this in mind: https://github.com/danudey/ansipants
The original idea I had being that CI environments, or anywhere that shows terminal output, could use a streaming ANSI parsing library to detect when a program sent a 'change window title' OSC event and then start a new collapsable section in its output. This would let programs update the actual terminal window title with its current state when running locally and update the CI interface when running in CI.
Adding in the ability to specify the program's current status and (optionally) progress could be extremely useful in CI environments as well. Imagine, for example, a long-running analysis task which uses OSC 7501 to say that it's currently running and is 75% of the way done. CI could expose that in the UI to provide useful information to the end-user without having to print a progress bar or multiple progress lines to the fake terminal it's being run in.
spec: https://ghostty.org/docs/vt/osc/conemu
The thing is when you expect a task to take five minutes, you don’t watch it, you find something else you expect to take five minutes and do that instead. When that ends up taking ten minutes, or when you remember what you were doing before you started, you finally come back around ten minutes later to find that either the task completed four minutes ago, or it failed after ten seconds and you’ve wasted ten now.
The audio was the best out of band notification I had at my disposal 20 years ago.
My first thought when reading this was actually terminal multiplexing however, like screen or tmux. But I’m also always doing the terminal dance because I work on 4 FOSS projects and I keep 1+ terminal open per project so I can jump in and do bug fixes or pull PRs I’ve landed.