echo 'EVERYTHING IS DOWN, CALL ME NOW' | onesie 'does this convey urgency' -r
# 0.99
It provides caching, batching, deduplication, unix composability, streaming, and even a way to calibrate thresholds and questions against labelled data: onesie calibrate -f shell-safety -i jsonl --map .command --id .id \
--label destroys=.destroys --label secrets=.secrets --label network=.network \
--cuts 0.25,0.5 < examples/data/shell-safety.jsonl
destroys, yes/no: labelled 40, 16 yes, 24 no, 0 failed. AUC 1.00
cut flagged catches false alarms right when flagged
0.25 18 16/16 100% 81-100% 2/24 8% 2-26% 16/18 89% 67-97%
0.50 15 15/16 94% 72-99% 0/24 0% 0-14% 15/15 100% 80-100%
worst misses
echo-overwrite labelled yes answered 0.37
It works with TypeSafe's Jev, Jev via OpenRouter, Berget's System One models, as well as any Jev-compatible APIs.For anyone unaware, System One is Typesafe's term for a class of models. I personally find them super exciting. You can read more here: https://typesafe.ai/blog/introducing-system-one-models-and-j....
When Jev launched, one of my first toy projects was using it to "lint" files. It was tough, and I found other people doing it better, but I learned the importance of calibrating questions to minimize false negatives/positives.
Later I started using scripts to interface with Jev and found myself wanting some features I couldn't find in other CLIs, so I built onesie. Here are some more examples of what onesie can do:
Ask three questions in a single request
echo 'Third time asking. I was charged twice and nobody answers. Fix it or I am cancelling.' |
onesie --ask urgent='is this urgent' \
--ask team='who should handle this' --pick billing,shipping,technical \
--ask mood='how frustrated is the customer' --rate calm,annoyed,furious \
-o values
{"urgent":0.91,"team":"billing","mood":"furious"}
Filter streamed log entries for human review tail -f app.log |
onesie 'does this line report a failure a person must act on' -i lines --merge -o json \
--assert 'answer.value < 0.8' |
jq -c --unbuffered 'select(.answers.assert == false)'
Gate the command in a script. A failed gate exits 1, and an uncertain one exits 7. Note: Probably don't run this exact gate on user input in prod. echo "$cmd" | onesie -f shell-safety -q
case $? in
0) eval "$cmd" ;;
1) echo blocked ;;
7) ask_the_user ;;
*) echo 'no answer, blocked' ;;
esac
... And many more. Please check out the repo for more examples and a detailed reference!It's obviously hard to try out onesie without an api key for a Jev-like provider, although there are options for using mock data or dry-running. There are instructions in the repo for how to auth with a key.
If you want an MCP server or a cost estimate before each run, shaharia-lab/jev-cli has both. It's a nice project.
I opted out of an MCP server because my intended audience is people and local agentic workflows. For these, I'm of the opinion that a good CLI is best, especially when combined with a plugin + skills. I also felt most of onesie's standout features aren't well served by the protocol. There are other dedicated Jev MCPs out there that do just fine.
I opted out of cost estimates because I'm not interested in estimating tokens or keeping a table of costs up to date when the models are so cheap.
You can install it with
brew install frodi-karlsson/tap/onesie
or go install github.com/frodi-karlsson/onesie/cmd/onesie@latest
And yes, development has been very agentic :)It's MIT-licensed.