我不得不把才华埋葬在昨天
A WeChat official-account article under the byline "intlsy's Doghouse" spread widely across the global AI community on September 14, 2026. The author, Liu Sheng, writes in the first person that he is a DeepSeek operator engineer, claims to have personally written the main Attention operator for DeepSeek V4.1 Flash, and publicly judges that "in another six months or a year, the operators AI writes will most likely be on a par with mine, or even surpass me."
Liu Sheng is no ordinary developer: he is top-tier talent holding up DeepSeek's compute foundation (a member of Peking University's Turing Class, class of 2021, and former captain of the Weiming supercomputing team). He has just delivered the core operator for DeepSeek V4.1 (the main Attention operator) and took a deep part in upgrading the DeepEP V2 expert-parallel communication library. Yet this very man, who forged the AI engine with his own hands, having watched models break into his professional heartland, has to admit he is being replaced by AI and faces being forced to "change trades."
The most eye-catching passage, at the end of the article, portrays Dario, the founder of Anthropic, as Hitler, on the view that Anthropic holding AGI is the same as Hitler holding the nuclear bomb. For the first time, a professional from China stands on the moral high ground. And his peers — Anthropic, OpenAI — happen, in this light, to stand in the moral lowland. Reposting the article here in full.
=====================Below is the original text of Liu Sheng's article====================
A few days ago, DeepSeek v4.1 was released, pushing the height of small-model capability up another notch.
AI is developing far faster than anyone expected. From the original ChatGPT, which could only babble its way through chat with a context length of a few thousand tokens, to OpenAI o1, DeepSeek R1, and Kimi K1.5 Thinking with reasoning ability took a mere two short years; from reasoning models to today's agents that fluently execute commands in all kinds of harness tools and complete complex tasks took only a year and a half. It is hard to imagine what AI will be like another one, two, or three years from now — how powerful it will be, whether it will already possess the ability to self-evolve, and how deeply it will have penetrated fields such as embodied intelligence.
AI Is Getting Better and Better at Writing Operators
In my own field of operator design and writing, AI has advanced just as quickly. Within a single year, it has gone from a little helper that could only look up documentation, read code, and find bugs for me into an operator master that can independently read CUDA, PTX, and SASS code, use professional tools to analyze the stall time of every instruction, and then optimize operators on its own. I believe that in the near future it will also gain the ability to independently design operator schedules, evaluate the performance of different scheduling schemes, and implement and optimize them.
Of course I am proud of DeepSeek v4.1's success — after all, its main Attention operator was written by me [1], and its excellence is a form of affirmation of my operator. But the wheels of the era roll relentlessly on, and no one can hold back the advance of technology. I know full well that in another half year or a year, the operators AI writes will most likely be as good as the ones I write, and may even surpass them. AI can think 300 tokens in a second, type out a command line in half a second, and finish a piece of code in twenty seconds — I cannot. AI can keep improving along every dimension of model depth, thinking intensity, volume of tool calls (frequency of interacting with the environment), even parallelism — I cannot.
Humanity has never hesitated, since ancient times, when it comes to destroying itself. Knowing full well that "the better my operators are, the faster our new models will train and run inference, the faster our models' capabilities will improve, and the sooner I will be replaced" — why do I still choose to optimize operators with all my might? Part of it is that writing operators really is like playing a video game for me and gives me enormous pleasure. The moment I invent a new technique, or watch my operator's performance numbers climb, the thrill is no less than what a speedrunner feels on breaking their own record. And when I see my operator's performance far outstrip the vendor's official operators, a great sense of pride wells up in me. But beyond that, there is a more important reason: even if I simply "let it all rot" from here on, or even deliberately threw up roadblocks to slow down model training, everyone else's models would keep developing all the same and take me out without fail. "Of course I hope I will not be revolutionized, but if revolution there must be, I want the person revolutionizing me to be myself." When everyone is this determined to destroy themselves, I have no choice but to join this brutal arms race.
And What About Me?
When the day comes that AI truly writes operators better than I do, what will become of me?
My judgment: I will not be "out of a job," but I will have to "change trades." My rice bowl can still be saved, but this may mean I never again get to do the work I once loved.
I once made a judgment about the changing times and my own future position: because the times are changing far too fast (AI's development, described above, is a good example), I simply cannot foresee what will happen in five or ten years. But no matter what, I believe that with my breadth of vision, my judgment, my initiative, and my intelligence, I can stay seated at the era's card table and once again stand at the crest of its wave. This judgment, however, can only guarantee that I will not be "out of a job"; it cannot guarantee that I will not need to "change trades" — or rather, it is a judgment that encourages me to avoid unemployment by changing trades.
So what does changing trades mean? It means giving up the field of operator design, writing, and optimization that I have cultivated for years and love deeply, and turning instead to being a "mecha pilot" for Agents. Until now, my interests, the things I was good at, and what industry needed were basically aligned; now AI has turned what I am good at into something it is better at, and industry's demand has drifted from "people who can write high-performance operators" to "people who can use AI to produce high-performance operators faster." To fit industry's needs, I will inevitably have to abandon the direction I loved and turn toward an unknown new one. I believe I can keep producing operators at high quality and high efficiency on the strength of my understanding of engineering, of upper-level model needs, and of the underlying hardware, and I know I may come to love this new direction (or may not) — but the feeling of having your love taken away from you is genuinely hard to bear. That quiet joy of sitting at my desk and calmly writing operators for a whole afternoon may become a swan song this very summer. I have no choice but to bury my talent in yesterday and go become a mecha pilot. My hands have picked up some gears, but my heart has lost some of its rhythm.
A vivid analogy: suppose you have mastered the craft of knitting sweaters, especially good at knitting all kinds of patterns and pairing all sorts of colors. The sweaters you knit are solid in quality and beautiful in design, and wealthy people from every village for miles around come to ask you to knit for them, earning you quite a bit of money. At the same time, you deeply enjoy that feeling of sitting by the window, brewing a pot of clear tea, gazing at the green hills, flowing water, cattle and sheep, and curling cooking smoke outside, and quietly knitting away a whole afternoon. But one day someone invents a miraculous machine: provide yarn and a pattern, and it knits a sweater automatically, with quality and texture no worse than your handiwork and at a speed far beyond yours. You know perfectly well that with this machine your peers can easily reach the level you once had, so you have no choice but to use it too. You also know that with the knitting skills you built up over the past twenty years, even once everyone has a machine, the speed and quality of your knitting can still outdo your peers'. But that pleasure of listening to the rain by the window, threading needle through yarn, and letting the hours pass slowly is, in the end, crushed under the roar of the machine.
I know this feels helpless, but there is no way around it. The rice bowl can be kept, but the old love will most likely have to be given up. I am someone who keeps reason and emotion in fairly separate compartments: I can be very rational when a problem calls for reason, but at times an emotional side shows too. I remember bawling my eyes out when I moved out of the rented flat I had lived in for a year, unable to part with the memories there. Saying goodbye today to that earlier era of hand-written operators and human-brain optimization is undoubtedly crueler still.
I don't know whether any readers feel something similar, but I suppose this is simply how it has to be.
And What About People?
While AI keeps advancing, I also have worries about a few questions:
- Aren't today's students most likely more inclined to use AI to finish their assignments, especially the hands-on labs? Picture two choices in front of you: one is grinding miserably through eight hours on a lab and maybe not even getting full marks; the other is firing up an AI model and, for the cost of a few cents and a few minutes, having the AI write perfect-score code outright. Which would most students choose?
- The point above will leave huge numbers of students severely short on engineering ability — the ability to organize code, to build systems, to think ahead about future needs and design for them in advance, to abstract, and so on. So with AI growing ever more capable, is this "engineering ability" still necessary? Will these engineering abilities be gradually cast aside by the times like the old ability to "fluently write x86 assembly," or will they keep lasting value like the ability to "understand the entire computer system from software to system to hardware"? If the latter, we are in danger — a person with poor engineering ability, paired with AI, can produce mountains of shit code several times faster than before, planting all manner of hidden hazards in systems and making this world even more of a slapdash amateur operation.
- In the society of the future, will power matter more than technical skill or IQ?
These questions may have to be answered by the era itself.
Conclusion
As AI develops, the society of the future may drift toward two extremes: communism and Cyberpunk 2077. In the former, productivity is enormously liberated and people's living standards visibly rise (I'll stop there, or I'm afraid this won't get past the censors); in the latter, a handful of tech companies control most of the resources, and only a tiny few can use the most advanced AI and every manner of technology, achieving something close to "mechanical ascension," while most people can only lay their hands on very feeble AI. Crossing class lines will become harder and harder: you must first hold the strongest AI before you can cross a class — a kind of infinite loop takes shape.
Guess what happens if Anthropic holds the world's most advanced AI forever: will the society of the future turn into communism or 2077? Take a guess.
So I still believe the most cutting-edge intelligence should be supplied to everyone, in an open and cheap way. I do not trust Anthropic or OpenAI to do that, and I especially do not want Anthropic to hold the most advanced artificial intelligence or AGI — to put it in exaggerated terms, the seriousness is no less than letting Hitler get atomic bomb technology before the Allies. This is also why I chose, and have insisted on staying, at DeepSeek: we research AI that is powerful, fast, and benefits everyone, and we open-source it — perhaps that can pull the world back a little from the 2077 end.
May all be well in the world of the future. May all the beauty be blessed.
[1] "Main Attention" covers only the MQA attention with head dim = 512; it does not include the indexer used to select the top-k important tokens — that part was written by other colleagues (of similarly very high skill) and their AI Agents.
Original article link:
https://mp.weixin.qq.com/s?__biz=MzcwMjI0Mjc0OQ==&mid=2247483692&idx=1&sn=534e145b7b97f9094a721cbf41dde346&chksm=f5b6102c3e7ba963f331ef4de9b68e3daaf6e13036d84be2ab6b27254b13d4c0ad31cac52bad&mpshare=1&scene=1&srcid=0915tDV7i8bDzGI5AiUM1NuG&sharer_shareinfo=df8791df4fb10516361cfb2e371f274e&sharer_shareinfo_first=ffc94c3a3d63b69df17381a9c8f6a1f8#rd