9 comments

  • atkrista 16 minutes ago
    I would just LOVE to see all the behind-the-scenes shithousery both companies are employing to one-up the other in this, largely, 2-horse AGI race. Someone should make a mockumentary when all is said and done!
    • nbardy 4 minutes ago
      I think it's weirdly just a choice of deciding to cut releases.

      We already know OpenAI has "bel" that is MUCH better than astra and is being used internally

    • petesergeant 7 minutes ago
      > largely 2-horse

      The absolute frontier is largely 2-horse, but the rest of the pack is very close behind, which I'm grateful for. Grok, Facebook, and the Chinese vendors are producing excellent models.

      • bayindirh 1 minute ago
        Gemini is also pretty nice for researching things. It turns out that having the whole indexed and having unlimited access to YouTube is a force multiplier of some kind.

        Since Google has their own TPUs, TPS is also pretty high w.r.t. Claude, for example.

      • Bluestein 2 minutes ago
        ... and, must be said a plethora of largely unsung, small, unknown "labs", outfits, "researchers" and the like. There is a long tail of smart people having at this. I guess, sheer compute aside, I think much progress - or, at least, important pieces thereof, will come from there.-
        • bayindirh 0 minutes ago
          There are some niche research areas where bog standard machine learning algorithms make miracles. LLM is just the poster child. AI/ML is a much larger and wider research area.
    • TeMPOraL 14 minutes ago
      AGI will make one, about humanity, after we're all gone - "They were so dumb, they just deserved to die".
      • cindyllm 11 minutes ago
        [dead]
      • f6v 7 minutes ago
        The sooner the better, brother.
  • madsgarff 11 minutes ago
    It would be very nice if artificialintelligence.ai actually, from a UX perspective, did the models in more than one thinking mode. I use claude, and I wanna build a feeling for what high, medium, etc. actually gives me. So far their comparisons, and having tried several different models for my work, has given me a feel of what 50 intelligence actually is. And I believe it would be be of even greater value to get a feel inside the single model I actually use, as most people do, because not many, I believe, switch heavily between models when working. I understand that the cost here is greater but the model provivders should obviously give you free access, because of the great work you are doing.
    • nextaccountic 8 minutes ago
      > It would be very nice if artificialintelligence.ai actually, from a UX perspective, did the models in more than one thinking mode

      But they do. See for example the pareto curve they have, try to locate GPT-6 Luna (max), (xhigh), (high), (medium), (low)

    • girvo 0 minutes ago
      It does!

      …but not for all models, which is pretty annoying.

  • egeozcan 18 minutes ago
    Worries of AI going rogue take so much attention that no governance body seems to care about the shady subscriptions and limits business.
    • semiquaver 8 minutes ago
      Do you mean the subsidized subscriptions which allow individuals to pay a tenth of the normal API cost for tokens?
  • max979 4 minutes ago
    Wild pace. Guess they found a critical bug or a quick win to push it out so fast. Astra is a high bar.
  • Pythagon 4 minutes ago
    Is this a duplicate thread of this? https://news.ycombinator.com/item?id=49896586
  • moomin 21 minutes ago
    This has got to be a panic move from OpenAI, right? They’ve had some bad press lately because from their billing changes, and Anthropic have finally released a fast, relatively cheap Opus with improved written English.
    • nba456_ 20 minutes ago
      Didn't they announce the billing changes the same day?
  • Dinuda 59 minutes ago
    After 5.5, I basically don't notice a jump in model performance, other than my usage ending sooner.
    • user43928 0 minutes ago
      I do.

      The results are less buggy, animations are much better.

      It can work autonomously for hours and the result is decent most of the time.

      That wasn't usually the case with 5.5, which needed more feedback and iterations to get things right.

    • zero1009 21 minutes ago
      Felt the same until I started using Luna. I feel like I get similar performance, but faster, and my usage lasts so much longer.
    • tom1337 18 minutes ago
      Kinda same but I miss my 5.3 Codex. Thing lasted forever on my $20 subscription and with detailed prompts was able to pretty much implement everything I requested it to do with a acceptable quality.
    • ModernMech 4 minutes ago
      Same. 5.5 got work done then 5.6 was also fine then 6 was maybe not quite as good. Now with 6.1 they are cutting usage and raising prices and introducing ultra fast mode, but things were good enough 5 months ago.
    • ndbe 19 minutes ago
      [dead]
  • 77rushi77 6 minutes ago
    [flagged]
  • tobin1994 15 minutes ago
    [dead]