Kimi K3-256k

(kimi.com)

110 points | by monneyboi 51 minutes ago

8 comments

  • illithid0 10 minutes ago
    This was posted 38 minutes ago, and as of 20 minutes ago, several Anthropic services are now designated as having a "major outage".

    Doubt these are related, but it made me laugh a little.

  • hawtads 32 minutes ago
    This is just an API level change right? The model itself should be the same I think.
  • dgritsko 24 minutes ago
    This isn't quantized, right? Just a smaller context?
    • DSingularity 14 minutes ago
      Its 256k context window. Quantization is orthogonal. We cant really tell directly so it could be quantized.
  • wxw 24 minutes ago
    > k3-256k is now available. Within 256k context, it delivers the same results. k3 (1M) consumes about twice as much quota as k3-256k.
  • madihaa 39 minutes ago
    That's actually nice! I usually try to stay below 200k context anyway.
    • giancarlostoro 35 minutes ago
      For me the sweet spot is somewhere under 500k depending on how extensive I want to get. You can build up a sizable effort project in half a million tokens with Claude, with Claude having all the context from ground 0 to wherever you're off at.
      • cyanydeez 30 minutes ago
        I'm always curious what you guys are working on; every git repo I've run a local model on and stick below <100k to increase speed seems effective enough to scope patches and changes.
        • KronisLV 6 minutes ago
          My current Claude Code session has been going on for like 35 hours and has used up around 400 million tokens, thankfully almost all of those being cached (95-98%) - pretty typical for long form agentic work.

          First you spend like 2-3 hours working on a plan, once you have that you just tell the model to go and implement it, do adversarial sub-agent review loops before each commit and also make sure that all tooling and tests pass (including coverage requirements). You do need to poke it in a slightly different direction every few hours, though. Not even any novel work, just some refactoring and SSE notification hardening, bug fixes, alongside environment tuning and getting rid of some bottlenecks (also migrated from Oracle to PostgreSQL but that's mostly done).

          That said, Kimi somehow manages to use less context in the main thread than Anthropic's models (even when you use sub-agents and also dynamic workflows in Claude Code), might have something to do with either how the model is tuned or their Kimi Code harness - because even in most of the longer form sessions it doesn't seem to fill up quite as quickly (note: because the kimi vis tool doesn't have a full summary view across all agents, these are the main long running agent stats across some sessions, not sub-agents):

            total tokens    cache hit rate    wall time    peak context
            283M            98%               3963m        466k
            258M            97%               2724m        467k
            98M             94%               1353m        393k
            67M             97%               614m         434k
            75M             98%               1447m        498k
            53M             99%               191m         375k
            6M              96%               139m         124k
            7M              98%               86m          118k
            11M             99%               61m          147k
          
          I could see 256k context being sufficient for all sorts of work, even if intermediate progress/plan tracking files and docs might have to be used along the way, in addition to whatever plan support the harness has (for example, if you document something that will be relevant for load testing you might need that in 10 turns but not during the ones before then).
        • jdoe1337halo 9 minutes ago
          They are just talking to the model in CC, while staying in a single thread. Doubt they have any actual coding knowledge to compartmentalize different problems in the codebase.
  • lukan 11 minutes ago
    Since Claude is the first time for me really, really out (TIL against my wished about https://status.claude.com/), I am now interested enough to see what else works. But ... when I click pricing, I see "Join a waitlist". Wtf? Are they really that good, so were totally surprised and overwhelmed by the requests, is this a marketing stunt, or do they just don't have the hardware being in china?
  • ibuildproducts 21 minutes ago
    omg! new model!!
  • superloika 38 minutes ago
    The bells tolled today, but nobody came to church.