For engineers
Try SAVI on your laptop, in minutes.
Add one wrapper to the AI client you already use. No account and no network calls: every call prints in your own terminal, and you get one report file you can send to anyone.
1Install it
pip install "savi-sdk[openai]"
2Swap one line, and switch on the report
import os
from savi import SaviOpenAI, local_report
local_report.start() # readable lines in your terminal, and a report file at the end
client = SaviOpenAI(api_key=os.environ["OPENAI_API_KEY"], local_mode=True)
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Summarise this contract."}],
)3Run your code and read your terminal
[ OK ] openai gpt-4o-mini-2024-07-18 in 214 out 96 812 ms cost ~$0.000090 [CACHE] openai gpt-4o-mini-2024-07-18 in 214 out 96 14 ms cost $0.00 answered from cache [ OK ] openai gpt-4o-mini-2024-07-18 in 330 out 60 701 ms cost ~$0.000085 personal data: EMAIL_ADDRESS x1 [FAIL ] openai gpt-4o in 0 out 0 30.2 s cost - failed timeout Summary (this computer only; nothing was sent anywhere) 8 calls, 1 failed tokens in 3,842 / out 978 average response 1.1 s cost about $0.01. Cost covers 6 of 7 calls. No price for: mistral-small-latest. 1 answered from the cache, saving about $0.000090
4Open the report
When your program ends, SAVI writes savi-report.html. It is one file. Open it in a browser, or send it to a colleague. It holds counts, models, tokens, times and costs. It never holds what was asked or answered.
How to read a line
- OK, CACHE, FAIL: whether the call worked, came from your own cache, or failed.
- in and out: the tokens each call used, and how long it took, from your code's side.
- cost with a ~: an estimate from the prices you enter. A model with no price is left out of the total, and the summary names it.
- personal data: the types found, never the text itself, once you add the personal-data option below.
- Repeats: the report notices when the same prompt goes out again and again. That is how a looping agent gives itself away.
Sample output from the SDK's demo command, shortened. The calls are made up. Streamed calls (stream=True) are counted too.
Start on your machine. Grow into the company record.
Local mode, free
What you see on your laptop
- Tokens in and out for every call, streamed ones included
- How long each call took, and which ones failed
- Repeated prompts, and answers served from your own cache
- A dollar cost, from your own prices in a small file
- One report file you can send to anyone
- Flags when a prompt contains personal data, with the personal-data option installed
The personal-data option is a larger install (pip install "savi-sdk[pii]"plus a language model download). Nothing leaves your machine either way. If something does not work, run python -m savi.doctor for a setup check that never prints your keys.
With a SAVI account
What your whole company gets
The same calls join one record across every team and tool, with spend by team, rules you can try before enforcing, and a signed audit trail. Your code stays the same: you add your SAVI key and the wrapper sends the signals instead of printing them.
Wrap the client you already use
Each wrapper sits around the provider's own client, so your calls and your prompts do not change.
SaviOpenAISaviAnthropicSaviAzureOpenAISaviVertexAISaviBedrockRuntimeSaviCohereSaviMistralOpen source
The SDK is MIT licensed, on PyPI and GitHub. Read every line that touches your calls. The SAVI service it reports to is hosted.
What is not covered yet
Agent frameworks such as LangChain, LlamaIndex and CrewAI, and gateway setups, are not officially supported. They may work, but check with us before you rely on them.
Want your whole company on one record?
Try it on one laptop first. When you are ready for every team, we will set it up with you.
