thank you for this! It reminds me of "What Uses More" website by Jon Ipolito at University of Maine--if you're not familiar with his work, he might be an interesting collaborator because you're both trying the be thoughtful about this work!
I'm unsure of your numbers for "coding agent session." I'm often burning hundreds of thousands of tokens for a single pull request. 22k feels like massive lowballing
I can’t find a single good number to average around and it seems to depend a lot on what specific people do, so I’m just leaving them to multiply it by the number of prompts
It's been running for about 24 hrs now and it's probably about halfway done. This is by far the longest one I've had, though.
The prompt is basically "ultracode Find all the bugs in our repository but only run one subagent in parallel because I want to use my quota for other stuff too".
Wow. Assuming these numbers are correct, this changes my mind on carbon use and I'm not sure how I'm going to reconcile my dissonance. If you could pick the following apart and tell me how it's wrong, I would be grateful.
I very much enjoy using coding agents. It's almost addictive how fun it is to use them. In the past week, Fable 5 has been in a limited trial period so I've been fully using my $20/month plan for Claude Code. This has amounted to ~500k Fable 5 tokens per day according to the Claude Code dashboard.
I've been considering upgrading to Claude Max, which is 20x the tokens. That would be roughly 10 million tokens a day. There are many people on the internets buying multiple Max subscriptions so they can get more Fable 5 tokens.
If I were spending 10 million Opus 4.8 tokens(Fable isn't an option in the calculator) a day I would have to do all of the following to offset that usage (none of which I currently do):
- Live car-free
- Switch to green heating
- A deep retrofit of my home
- Go vegan
I would love to know how Fable changes these estimates. Coding agents seem like the next killer app for these models and Fable is a leap forward on their usefulness. The better the models get the more people will want to use them, especially at scale in business. Anthropic claim more powerful models are coming in a few months, presummaly more carbon per token as well?
Like the personal carbon footprint calculators of yesteryears, I feel like this misses the point. The average and median user of AI is not using anywhere near enough AI to make a meaningful blip on environmental costs, in much the same way the average or median’s person’s carbon footprint is a minuscule contributor to global warming. With both issues, there are some “heavy polluters” who need to be addressed.
AI data centers are being rushed as the AI world is in a race for compute. They are being forced through permitting and questions about whether local infrastructure can support them are being deferred. Even if in the grand scheme of things the water or energy use of data centers appear reasonable, whether local infrastructure can support it is a different issue. This is especially relevant in the USA, which has neglected its energy grid and water management responsibilities for decades.
On the consumption side, a handful of individuals are consuming the vast quantity of AI tokens. (100s of billions of tokens vs 10s of thousands). The fad from earlier this year of token maxing has run its course, and I think we are going to see more discussion about ROI of AI solutions. See https://www.notboring.co/p/return-on-tokens-rot
I agree, but I can only talk about one thing at a time, and a lot of people are under the impression that their AI use adds a lot to their emissions and water use
Fair. Have you found them receptive to your messaging? I get the feeling many people who believe this stuff feel existentially threatened by AI. It sort of reminds of the free trade debates of the 90s and early 2000s. Their economic arguments were terrible but reflected real economic insecurity / vulnerability.
If you look at Anthropic's latest mechinterp papers (which I think are all nonsense in their claims, but that aside) they are clearly describing the operation of a dense model. Their J-Space one in particular would need to do something to deal with the routing function if there were one.
Never in any of their mechinterp papers do they mention having to adjust their methodology for a routing function, or look at different sub-networks. Not once. It is conspicuous by its absence.
It is, ultimately, an ideological choice: clearer mechinterp experiments over model energy efficiency.
Also, it's a very reasonable inference from the fact that they haven't trained it multimodally. They train it on text, and then bolt on a vision encoder that converts images to text embeddings. Why would they avoid doing that, unless they were desperate to keep parameter count down?
Its strangely small token vocabulary could also be a way of keeping parameter count down, though there could just be other reasons for that.
Do you have any plans or ideas for how to add equally solid data for image/video generation (as examples) to the visualization (while we wait/hope for EcoLogits to expand their data set)
The calculation for daily overall water use doesn't seem to reflect any lifestyle changes except for where the user lives (the carbon calculator works fine in this regard).
Nice one, thanks for sharing. Rather than cutting my AI use, I'm going to start microwaving for 3 minutes less each day.
thank you for this! It reminds me of "What Uses More" website by Jon Ipolito at University of Maine--if you're not familiar with his work, he might be an interesting collaborator because you're both trying the be thoughtful about this work!
Right on. Thanks for this, and for the eminently sensible suggestion you make to move the debate to things that actually move the needle.
I'm unsure of your numbers for "coding agent session." I'm often burning hundreds of thousands of tokens for a single pull request. 22k feels like massive lowballing
I can’t find a single good number to average around and it seems to depend a lot on what specific people do, so I’m just leaving them to multiply it by the number of prompts
I’ll put it back to 100k to be safe
It's very context dependent but average is probably growing fast.
Here are stats from my current Claude Code session. Single prompt except I had to switch from Fable to Opus in the middle of it.
claude-fable-5: 134.4k input, 687.2k output, 92.8m cache read, 3.3m cache write ($170.25)
claude-opus-4-8: 1.2m input, 3.0m output, 305.5m cache read, 12.0m cache write ($309.11)
It's been running for about 24 hrs now and it's probably about halfway done. This is by far the longest one I've had, though.
The prompt is basically "ultracode Find all the bugs in our repository but only run one subagent in parallel because I want to use my quota for other stuff too".
Wow. Assuming these numbers are correct, this changes my mind on carbon use and I'm not sure how I'm going to reconcile my dissonance. If you could pick the following apart and tell me how it's wrong, I would be grateful.
I very much enjoy using coding agents. It's almost addictive how fun it is to use them. In the past week, Fable 5 has been in a limited trial period so I've been fully using my $20/month plan for Claude Code. This has amounted to ~500k Fable 5 tokens per day according to the Claude Code dashboard.
I've been considering upgrading to Claude Max, which is 20x the tokens. That would be roughly 10 million tokens a day. There are many people on the internets buying multiple Max subscriptions so they can get more Fable 5 tokens.
If I were spending 10 million Opus 4.8 tokens(Fable isn't an option in the calculator) a day I would have to do all of the following to offset that usage (none of which I currently do):
- Live car-free
- Switch to green heating
- A deep retrofit of my home
- Go vegan
I would love to know how Fable changes these estimates. Coding agents seem like the next killer app for these models and Fable is a leap forward on their usefulness. The better the models get the more people will want to use them, especially at scale in business. Anthropic claim more powerful models are coming in a few months, presummaly more carbon per token as well?
It’s possible the numbers are off, but yeah this could also just actually add up to being a significant part of your emissions!
Thank you! This is wonderful
Like the personal carbon footprint calculators of yesteryears, I feel like this misses the point. The average and median user of AI is not using anywhere near enough AI to make a meaningful blip on environmental costs, in much the same way the average or median’s person’s carbon footprint is a minuscule contributor to global warming. With both issues, there are some “heavy polluters” who need to be addressed.
AI data centers are being rushed as the AI world is in a race for compute. They are being forced through permitting and questions about whether local infrastructure can support them are being deferred. Even if in the grand scheme of things the water or energy use of data centers appear reasonable, whether local infrastructure can support it is a different issue. This is especially relevant in the USA, which has neglected its energy grid and water management responsibilities for decades.
On the consumption side, a handful of individuals are consuming the vast quantity of AI tokens. (100s of billions of tokens vs 10s of thousands). The fad from earlier this year of token maxing has run its course, and I think we are going to see more discussion about ROI of AI solutions. See https://www.notboring.co/p/return-on-tokens-rot
I agree, but I can only talk about one thing at a time, and a lot of people are under the impression that their AI use adds a lot to their emissions and water use
Fair. Have you found them receptive to your messaging? I get the feeling many people who believe this stuff feel existentially threatened by AI. It sort of reminds of the free trade debates of the 90s and early 2000s. Their economic arguments were terrible but reflected real economic insecurity / vulnerability.
Update: EcoLogits has their own calculator now at https://calculator.ecologits.ai/
There's no way that Claude uses less energy than GPT and Gemini. Claude is dense, GPT and Gemini are MoE.
Where are you getting the claim that Claude is dense? I can't find a source to back that up.
You also won't find a source that says Claude is MoE, since Anthropic hasn't said it is. If it were then I'm sure they'd boast about it like the other big vendors did when they got MoE working. But, here: https://www.digitalapplied.com/blog/moe-architecture-comparison-gpt-claude-deepseek-qwen
If you look at Anthropic's latest mechinterp papers (which I think are all nonsense in their claims, but that aside) they are clearly describing the operation of a dense model. Their J-Space one in particular would need to do something to deal with the routing function if there were one.
Never in any of their mechinterp papers do they mention having to adjust their methodology for a routing function, or look at different sub-networks. Not once. It is conspicuous by its absence.
It is, ultimately, an ideological choice: clearer mechinterp experiments over model energy efficiency.
Also, it's a very reasonable inference from the fact that they haven't trained it multimodally. They train it on text, and then bolt on a vision encoder that converts images to text embeddings. Why would they avoid doing that, unless they were desperate to keep parameter count down?
Its strangely small token vocabulary could also be a way of keeping parameter count down, though there could just be other reasons for that.
Awesome, I created a Swedish version here https://resurskollen.jardenberg.se
Thanks for sharing the code.
Do you have any plans or ideas for how to add equally solid data for image/video generation (as examples) to the visualization (while we wait/hope for EcoLogits to expand their data set)
The calculation for daily overall water use doesn't seem to reflect any lifestyle changes except for where the user lives (the carbon calculator works fine in this regard).