I tried the Fugu models with some real world tales in C# and unity using mcp and open code. I exhausted the $20 plan 5 hour window in one prompt to review my theme system and plan some color changes. So I upgraded to the $100 to see the implementation and result. Well the result was worse than Opus, incredibly slow, and I ended up exhausting the new 5 hour window and have used 35% of the weekly now and it hardly created something opus was able to do at a fraction of the time and cost.
Do what you wish with this info, but it seems to be a complete waste of $$.
Because Fugu is not an independent model. They just use multiple existing SaaS models from OpenAI, Anthropic etc in background, gather response and generate results based on these response.
They claims that combining the results of multiple AI models and generate final result by using their in-house proprietary model improve the quality than using the single backend model.
It cause all sorts of doubts like: is their in-house model really exists? Is their in-house model really capable?
Personally, even if their claim is correct, such feature can be easily implemented in client side like Claude Code etc, using the equally capable model from background model to generate the final result.
We provide a similar service for Godot instead of Unity, and 20$ plan being exhausted in one prompt on a top model like Opus sounds about right. That's the life when you pay API prices and can't afford 10x subsidies.
Sure, I understand the subsidization. Their limits are practically unusable and the marketing of "Focused working sessions for regular coding, review, research, and analysis throughout the week." is pretty disingenuous then.
I don't think a different way of paying should be considered a subsidy. From what I understand AI companies are making money on those plans.
The limits feels really usable to me, but maybe because I've learned to work within their limitations. It's like, it's always possible to consume more tokens for maybe more output.
Is this true even for more targeted prompts that aren't about looking over an entire codebase or w/e? I just stick to my sub pricing and find good success on targeted requests, but I wonder if I would run up against things even then if not for subsidies.
I tested Fable through Cursor; asked for ideas on how to make a data website I have less "Claude-like" (IYKYK what are the usual tells), and it spun out the most useless, Claude-like CSS styling ever, wasting $40 in 10 minutes.
The website was created through Opus, so you could also say the results were worse than Opus. (This is just to say that I had the same experience using the US models, so perhaps those Asian models are Mythos-like lol)
I tested Fable for a whole day. And my experience was quite the opposite. I was blown away. Admittedly, I did not try it through a middleman like Cursor. I used Claude Code CLI.
Me too. It’s output was fabulous. And it acted like a senior engineer - actually coding up hypotheses, testing them, finding problems and presenting good, usable recommendations backed by solid evidence and wisdom. It can probably do most of my job, which gave me a bit of an existential crisis.
I’ve paused my Claude subscription until they bring it back. Opus makes mistakes constantly, on every level of abstraction. Every time I look closely at its work I find problems.
If you actually give it an example of the style it can copy it well. even just screenshots of other websites or UIs. It just sucks at producing it itself.
Well, in part because the phenomenon has been discussed on Web forums that (a) have at this point made their way back into training data and (b) are accessible in Web searches that the model can invoke. And in part because the model can "know" what its initial instinct is and "decide" to go against it.
I experienced the same exact thing, however, I will say that I had misconfigured it on `pi` at first.
I was using its chat endpoint not the responses with tool calling and everything, and I haven't tried it again since, learned that recently and I am planning on giving it another shot.
This is useful info. For the couple of days that Fable was live - it was clearly a step above Opus 4.8 and I was able to get 8-10 prompts in using my $20 plan.
Do what you wish with this info, but it seems to be a complete waste of $$.