AI Industry
The Terms on Which Liang Wenfeng Took the Money
DeepSeek's first outside round, about 51 billion yuan in June 2026, reportedly put commercial backers in a partnership with a five-year lock-up and no votes, with China's national AI industry investment fund the one reported exception. A reported Shanghai listing would put those terms in public documents.
Tips, corrections, or questions? support@omniscient.media

[Editor's Note: This piece was researched, drafted and checked with the help of AI models made by Anthropic and OpenAI, two of the companies whose accusations against DeepSeek are described below.]
On April 28th, 2026, a Chinese securities wire reported a change in a Hangzhou company registry. The registered capital of the company behind DeepSeek had risen from 10 million yuan to 15 million, and its founder, Liang Wenfeng, had lifted his direct stake from 1% to 34%.[1] The same report noted that DeepSeek was said to be preparing its first outside fundraising.[1] Seven weeks later, on June 16th, that round closed at about 51 billion yuan, according to registry data reported by the tech site IT之家, and Liang himself was on the list of investors, beside Tencent, CATL, JD and NetEase.[2] Reuters had earlier reported that he was committing 20 billion yuan of his own money.[3]
The terms came with a caveat. According to The Information, in a report Reuters said it "could not immediately verify," outside investors were required to put their money into "a limited partnership managed by DeepSeek CEO Liang Wenfeng" rather than into DeepSeek itself, accept a five-year lock-up, and take no voting rights.[4] China's National Artificial Intelligence Industry Investment Fund was "the only exception", investing directly, with a vote and without the lock-up.[4] What the register shows independently is the shape: in July, registry data reported by the 21st Century Business Herald listed two new shareholders in DeepSeek, a partnership called 杭州程砺 that had taken in Tencent-, CATL- and JD-linked investors, and the national fund itself.[5] The same filing showed Liang leaving 杭州程砺 as a partner,[5] which sits awkwardly beside the report that he manages it, and which no document we could read resolves. Nor did we find one that says how his 34% direct stake, his reported 20 billion yuan commitment and the partnership add up to control of the company.
According to Reuters, people close to Liang described the lock-up as a way to screen investors and keep out capital seeking a quick exit.[2] The proceeds are earmarked for infrastructure, research, employee equity and commercialization.[2] Valuations vary by who's counting: IT之家 put the company at close to 400 billion yuan at completion, and Reuters later described the round as done at about 450 billion yuan post-money.[2][6]
When Liang said in 2024 that DeepSeek's problem had never been money, a hedge fund's trading profits were still keeping the lab independent. By June it had taken a great deal of outside money, and the reported terms, a five-year lock-up with no votes, would keep out the investors with "exit needs" he'd complained about in 2023. The public record doesn't settle whether Liang, Beijing or a reluctant market set those terms, and the only investor reported to sit outside them was the state's own fund.
A fund that stopped taking money
Liang has given almost no interviews. In one of the two long ones, with Waves (暗涌), the venture-capital publication of the Chinese tech outlet 36Kr, in 2024, he described his childhood in a sentence. He grew up in the 1980s in a fifth-tier city in Guangdong, and "my father was a primary-school teacher."[7] (The translations from Chinese in this piece are our own.) He pushed back on the idea that his firm appeared from nowhere: "What outsiders see is High-Flyer after 2015, but in fact we've been at it for 16 years,"[7] which dates the work to about 2008, where High-Flyer's own timeline begins, with the founding team exploring fully automated trading from scratch.[8]
Waves reported in 2023 that while he was studying artificial intelligence at Zhejiang University he was already convinced that "AI will definitely change the world," a conviction that in 2008, it noted, few shared. After graduating, according to the same report, he passed up the usual programming job at a big company for a cheap rented flat in Chengdu, trying one application after another, and failing, before settling on finance. An early friend who tried to recruit him, then building aircraft in a Shenzhen urban village, went on to found the drone maker DJI.[9]
High-Flyer, the quantitative hedge fund he founded in 2015,[2] tells its own history on its website. In October 2016, it says, the first stock position generated by a deep-learning model went into live trading.[8] It then built two computing clusters it calls Fire-Flyer, and a 2024 paper co-authored by DeepSeek researchers describes the second as 10,000 PCIe A100 GPUs,[10] bought, as SemiAnalysis later noted, before export restrictions on such chips.[11] That cluster is the foundation DeepSeek was later built on, and it was paid for out of the fund's trading profits, which IT之家 credits with keeping DeepSeek independent for years.[2]
The fund's run had a hard stretch. On November 15th, 2021, High-Flyer suspended subscriptions to all of its products, and an insider said it would stop growing and concentrate on improving its strategies at its existing size.[12] After record drawdowns that autumn, it told its investors that its AI had picked stocks that were sound on long-term value, but that "on the timing of buying and selling, it really did not do well."[12] By May 2023, Liang said the fund "basically doesn't raise money from outside much anymore."[9]
The argument he made
Liang's two long interviews with Waves, in May 2023 and July 2024, raised questions the 2026 cap table answers. In 2023 he stated the goal plainly: "What we want to build is general artificial intelligence, AGI."[9] The lab would do research and exploration and would "not do vertical products or applications."[9]
Money was the first obstacle the interviewer raised: without two or three hundred million dollars, you can't even get a seat at the table, so how would DeepSeek pay for it?[9] "We're also talking with different funders," he said, before explaining why that was hard: many venture investors "have reservations about research; they have exit needs."[9] SemiAnalysis's later account of the period points to a different cause for the missing outside money: High-Flyer funded DeepSeek itself because "outside investors had little interest in AI at the time, with the lack of a business model being the main concern."[11] Asked for the commercial logic, he said that if you insisted on one "you probably won't find it, because it doesn't pay."[9]
In July 2024 he gave the line that matters most for the 2026 round: "No fundraising plans in the short term. The problem we face has never been money; it's the embargo on high-end chips."[7] "In the short term" is a bounded promise, and read beside the 2023 interview, where he was already in talks with funders, it describes timing. He was describing a compute problem, with US export controls as the barrier, and the later fight over H200 licenses and DeepSeek's move onto Huawei's Ascend chips (both covered below) are consistent with that reading.
The rest of that interview describes the lab the new money went into. DeepSeek had started a price war in China's model market, and he called it an accident: "We didn't set out to be a catfish; we just accidentally became one,"[7] the catfish being the fish that keeps a tank of sardines moving. It had simply priced from its costs: "our principle is not to subsidize, and not to make windfall profits either."[7] On openness he said that "facing disruptive technology, a closed-source moat is temporary," and "we won't go closed source."[7] On the gap with the United States: "we often say Chinese AI is a year or two behind the US, but the real gap is between originality and imitation." More investment, he added, "doesn't necessarily produce more innovation."[7]
Work at DeepSeek was "entirely bottom-up," with roles that emerged rather than being assigned in advance, and anyone with an idea "can call on the training cluster's cards at any time without approval."[7] SemiAnalysis, writing in January 2025, described DeepSeek job ads that offered access to tens of thousands of GPUs "with no usage limitations."[11] Asked about Anthropic co-founder Jack Clark's description of the DeepSeek team, which the interviewer rendered as a group of "inscrutable wizards," Liang said: "There are no inscrutable wizards," only fresh graduates from top universities, fourth- and fifth-year PhD students working as interns, and people a few years out of school.[7]
What he signs
Liang signs DeepSeek's research as its lead: last-listed author, the conventional position for a lab's research lead, on focused papers such as DeepSeekMoE,[13] one of the authors of Native Sparse Attention, which won a best-paper award at ACL 2025,[14] and corresponding author when the reasoning model R1 was published in Nature in September 2025.[15] The papers credit others with the inventions; DeepSeek's V2 report says that "Huazuo Gao and Wangding Zeng have made key innovations in the research of the MLA architecture,"[16] the attention design that made DeepSeek's models cheap to run.
They're careful about cost, too. The V3 technical report said that "our total training costs amount to only $5.576M," and in the very next sentence that the figure covered only the official training run and excluded "the costs associated with prior research and ablation experiments."[17] The $5.6 million figure became shorthand for DeepSeek in January 2025 even though the paper had already listed what it left out. The most widely cited counter-figure, SemiAnalysis's estimate of about $1.6 billion in "the total server CapEx for DeepSeek," measures something else: a fleet of GPUs "shared between High-Flyer and DeepSeek."[11]
For the new shareholders, the 2023 commitments that matter most are the two that shape what they've bought into: open weights and no applications. "Open" has mostly held up as a description. R1's weights were released under the permissive MIT license in January 2025,[18] and later models followed. In August 2026 the first vision model in DeepSeek's V4 family launched through the company's API only, which our August 24th bulletin reported as the weights being withheld; ten days later they appeared on Hugging Face under the MIT license.[19] The April 2026 release note for the V4 series closed on a commitment to 长期主义, long-termism.[20]
The other 2023 commitment, no applications, is harder to square with what DeepSeek ships. DeepSeek runs a consumer chat service with its own privacy policy,[21] and in August 2026 it created a public repository for an agent framework, DeepSeek Harness, which it describes as "Everything is a Plugin."[22] In a transcript of a May meeting with investors, which circulated in late July and which DeepSeek has never acknowledged, Liang is quoted calling consumer and enterprise products "by-products on our road to AGI."[23]
The state's hand
How close Liang sits to Beijing is the question most often asked about him in Washington, and the public record answers it only partly. The formal contacts are few. On January 20th, 2025, the day DeepSeek released R1, he was among the entrepreneurs and experts who spoke at Premier Li Qiang's symposium on the draft Government Work Report; the official readout lists his name and records nothing he said.[24] At Xi Jinping's symposium with private companies on February 17th, 2025, the readout names six business leaders who spoke, among them Ren Zhengfei and Lei Jun, and Liang isn't one of them,[25] though a Chinese outlet describing state television footage reported that he attended.[26] We found no government office, title or honor held by him.
The informal pressure is better documented. In March 2025 The Information reported, as relayed by TechCrunch, that some DeepSeek employees had been "prevented from traveling abroad freely" and that the Chinese government was "playing a role in screening potential investors."[27] That was fifteen months before the first outside round closed. In May 2026 Bloomberg reported that the travel controls were being extended to top researchers at other private AI firms, some of whom were asked to surrender their passports to their employers.[28]
Set beside the cap table, that record puts the state at two points in the financing: a reported role in screening investors before anyone had invested, and the only outside shareholding reported to carry a vote.[4] None of this shows the government directing the company's research or products, and nothing we read claims it does.
The charges
The charges matter to the financing because two of them run through governments: the joint US advisory described below says the extraction it alleges likely happened with Chinese government awareness, and Washington has reportedly held off adding DeepSeek to its Entity List. That's exposure DeepSeek's new investors now share, and a listing would put it in front of public ones.
A note on who is making them. Several of the accusations below come from Anthropic and OpenAI, which compete with DeepSeek for the same customers and have a commercial interest in how it's regulated. Anthropic's September report also says DeepSeek tagged people using coding tools, including Anthropic's own Claude Code, and relayed some of their requests to Claude.[29]
On chips, the accusations are serious and unproven. The link between DeepSeek and a Singapore fraud case over servers that may have held Nvidia chips came from local media reporting "without identifying its source."[30] A senior State Department official told Reuters in June 2025 that DeepSeek had sought to evade export controls, but "declined to say if DeepSeek had successfully evaded export controls."[31] A year later Reuters reported that DeepSeek had been "approved by an interagency committee last year" for the Commerce Department's Entity List, but that the administration "has held off adding" it and others to avoid escalating tensions with Beijing.[32] We found no later report that it had been added.
Distillation, training one model on another's outputs, is where the charges are most specific. In a February 2026 memo to the House China committee, OpenAI said it had observed "accounts associated with DeepSeek employees" working around its access controls.[33] Anthropic reported the same month that it had traced a DeepSeek campaign of "over 150,000 exchanges" with Claude and could "trace these accounts to specific researchers," the smallest of the three Chinese labs it named.[34] Its September report put the figure attributable to DeepSeek at "over 12.1 million exchanges observed" in 14 days of July 2026.[29] We found no underlying data published by either company.
On September 8th, 2026, the NSA, CISA and the FBI issued a joint advisory naming DeepSeek first among six Chinese companies that, "likely with Chinese government awareness," had extracted outputs from American models.[35] It called DeepSeek's cost figure misleading because it "does not include the true cost of the data" obtained this way.[35] Its list has a problem of dates: it says DeepSeek distilled from models including Claude Sonnet 4.5 and Grok 4 "between late 2024 and mid-2025" to train R1 and V3,[35] but Grok 4 was released in July 2025 and Sonnet 4.5 in September 2025,[36][37] after R1 and V3 had shipped, and in Sonnet 4.5's case after the advisory's own window had closed. China's commerce ministry called the allegations baseless in fact and in law, and defended distillation as "in essence a neutral technical means" that American firms use too.[38] We found no on-the-record response from DeepSeek itself.
The distillation charges press on the distinction Liang drew most sharply. "The real gap is between originality and imitation," he said in 2024, and unless that changed, "China will always be a follower."[7] Distillation is a charge about training data, not about the architecture work the papers credit, such as MLA. But if the accusers' evidence holds up, DeepSeek was also leaning, at scale, on the kind of imitation he said China had to outgrow.
The censorship charge is the best measured, and it concerns the models themselves. In September 2025 the US government's Center for AI Standards and Innovation found that DeepSeek models echoed "4 times as many" inaccurate Chinese Communist Party narratives as American reference models on politically sensitive questions, testing weights downloaded directly, and that the censorship "occurs whether users interact with the model in English or Chinese."[39] DeepSeek's privacy policy says user data is stored in the People's Republic of China.[21] On safety, DeepSeek has signed China's voluntary Artificial Intelligence Safety Commitments, though OpenAI's memo notes that it "still has not published a clear safety framework or evidence of robust testing and independent red-teaming."[33] We found no statement by Liang on catastrophic risk in either interview or the leaked transcript.
The leak
The closest thing to Liang speaking at length in 2026 is a document DeepSeek has never acknowledged. In late July a transcript of a closed meeting with investors circulated on Chinese social media; an investor in DeepSeek told National Business Daily that the meeting took place in May and that the content was genuine.[23] The version mirrored on GitHub describes itself as a machine transcription tidied by AI that doesn't distinguish speakers.[40] Bloomberg said it hadn't "verified the authenticity" of the posts.[41] The only relevant public line from DeepSeek we found predates the leak: its April release note asked readers to "rely only on our official accounts."[42]
With that caveat attached to every word, the transcript has him making a case about resources. "Talent is not the bottleneck; resources are the biggest bottleneck," he says.[43] Nvidia's CUDA software moat "is being dismantled rapidly," and Huawei's chips are a workable substitute at a price: "four Huawei cards equal one Nvidia card."[43]
The transcript also undercuts the idea that money had stopped mattering. "Even if we spent the whole 50 billion, we couldn't afford" to train at the scale of the largest models, he says,[23] and if the company could turn all its money into chips, "we would turn all the money into cards without hesitation."[44] According to Bloomberg, DeepSeek then told prospective investors it was pausing a second funding round, a suspension that "stemmed in part from Liang's frustration over online reports about his comments to investors."[41]
The next test
Reuters reported on July 15th that DeepSeek was planning a fresh round at a valuation of about 500 billion yuan ($74 billion) ahead of a potential listing in mainland China.[6] Ten days later came the pause described above. Bloomberg then reported on August 6th that DeepSeek had resumed the round, seeking close to $8 billion, with Monolith Management in talks to take part,[45] and the South China Morning Post reported on August 26th that it was "expected to close before the end of August,"[46] a date that came and went.
On September 9th, Reuters reported, citing two people with knowledge of the matter, that DeepSeek had tapped CITIC Securities to prepare an initial public offering on Shanghai's STAR Market.[47] Sina Tech reported the same engagement that day, citing a person who confirmed it, and added that the two sides had not yet signed a formal listing-tutoring agreement; we found no comment from DeepSeek.[48] On September 24th, Reuters relayed The Information's report, citing two people with direct knowledge, that DeepSeek's annualized revenue run rate had reached $1 billion, more than double its level a few months earlier, and that the second round was still open, now targeting 50 billion yuan at a valuation of 500 billion yuan "by the end of October."[49] A run rate annualizes a recent stretch rather than reporting booked revenue, and this one rests on anonymous sources, so it's the number a prospectus would have to replace.
A listing is where the terms get tested in public. STAR Market rules allow a company to list with weighted voting rights, but only if the arrangement is in place before the IPO; it "may not be set up in any way after listing,"[50] a single share's votes are capped at ten times an ordinary share's,[50] and the weighted shares must be held by a continuing director or an entity that director controls.[50] None of the sources we read says whether DeepSeek plans to use that structure, keep the partnership, or unwind it. What they do show is that the choice has to be made before the listing, in documents the exchange will publish.
The other test is chips, where Beijing's approvals now matter as well as Washington's export controls, and they have been announced more often than they have been used. Reuters reported on January 30th that China had approved DeepSeek's purchase of H200s "with regulatory conditions that are still being finalised."[51] By July 8th, Reuters, which had reported in March that Nvidia won Beijing's approval to sell the chips, was reporting that Washington had licensed about ten Chinese firms to buy Nvidia's H200 while Chinese officials "have withheld approval so far."[52] The Information, in the report Reuters was carrying, said Chinese officials had told DeepSeek and others they "may soon receive permission to buy some H200 chips," in quantities that could total "fewer than 200,000."[52] The Financial Times reported on August 18th that small batches of H200s had since been allowed into mainland China, naming ByteDance and Tencent as recipients;[53] the reports we found don't say whether DeepSeek has received any. On July 7th, Reuters reported that DeepSeek was "developing its own AI chip, according to three people familiar," one "designed for inference."[54] The transcript DeepSeek has never acknowledged, from a May meeting that predates that report, quotes him saying that whether to make its own chips would depend on how much it paid off, and that "I hope we don't have to make chips."[23] A change of plan between May and July is one reading of the gap, and an imprecise transcript is another.
On September 29th and 30th, DeepSeek published code that moves some of its work onto Huawei's Ascend chips, including a matrix-multiplication library whose repository was created on September 29th.[55] Its statement said the two companies were "jointly advancing" a 128-chip supernode built on Huawei's Ascend 950,[56] and the documentation says the benchmarks were run on a "manually configured PoC HDK," a test setup, with Huawei's full-bandwidth commercial release planned for around October 15th.[57] The kernels are "built on TileLang,"[58] an open-source language created in 2024 by a separate group,[59] which DeepSeek says now carries most of the operator code used in training its V4 models,[56] though in the leaked transcript Liang is quoted calling it a technology "our house produced."[43]
So the next few weeks carry dated things to watch: the second round's end-of-October target; a listing-tutoring filing with China's securities regulator, which would start the clock on a STAR listing; Huawei's commercial release around October 15th, which DeepSeek's documentation says should give outside users the full-bandwidth configuration its own benchmarks relied on; and whatever DeepSeek publishes about its structure before it lists. If DeepSeek does file for a STAR listing, its prospectus would have to set out the full terms, including who holds the votes, beyond the shareholder names the registry already shows.
Sources
人民财讯 / Securities Times (证券时报网), via Sina (April 28, 2026), citing Qichacha registry data: 「注册资本由1000万元增加至1500万元」; 「持股占比由1%提高至34%」 Inline ↗
3 passages checked · read October 5, 2026
…缩小字体 放大字体 收藏 微博 微信 分享 0 腾讯QQ QQ空间 人民财讯4月28日电,企查查数据显示,杭州深度求索人工智能基础技术研究有限公司(简称“DeepSeek”)注册资本由1000万元增加至1500万元,其中梁文锋认缴的注册资本由10万元增加至510万元,持股占比由1%提高至34%。此前媒体报道称,DeepSeek可能正启动首次外部融资,腾讯控股和阿里集团或参与投资。 责任编辑:刘德宾 赛博对话 热搜时代 一天零一页…
…腾讯QQ QQ空间 人民财讯4月28日电,企查查数据显示,杭州深度求索人工智能基础技术研究有限公司(简称“DeepSeek”)注册资本由1000万元增加至1500万元,其中梁文锋认缴的注册资本由10万元增加至510万元,持股占比由1%提高至34%。此前媒体报道称,DeepSeek可能正启动首次外部融资,腾讯控股和阿里集团或参与投资。 责任编辑:刘德宾 赛博对话 热搜时代 一天零一页 无双 热浪之外 打印网页 --> 我要反馈 相关新闻…
…索人工智能基础技术研究有限公司(简称“DeepSeek”)注册资本由1000万元增加至1500万元,其中梁文锋认缴的注册资本由10万元增加至510万元,持股占比由1%提高至34%。此前媒体报道称,DeepSeek可能正启动首次外部融资,腾讯控股和阿里集团或参与投资。 责任编辑:刘德宾 赛博对话 热搜时代 一天零一页 无双 热浪之外 打印网页 --> 我要反馈 相关新闻 投资热点尽在新浪财经APP> 加载中 点击加载更多 阅读排行榜 评论排行榜…
IT之家 (June 17, 2026), citing Qichacha registry data and earlier Reuters reporting: 「整体融资规模约 510 亿元,企业估值近 4000 亿元」; 「此举旨在筛选投资者,排除追求快速退出的资本」 Inline ↗
5 passages checked · read October 5, 2026
DeepSeek 以 4000 亿元估值完成首轮外部融资:510 亿元到账,投资方含梁文锋、腾讯、宁德时代、京东、网易等 - IT之家 首页 IT圈 最会买 设置 日夜间 随系统 浅色 深色 主题色 黑色 投稿 订阅 RSS订阅 收藏IT之家 软媒应用 App客户端 要知App 软媒魔方 业界 手机 电脑 测评 视频 AI 苹果 iPhone 鸿蒙 软件 智车 数码 学院 游戏 直播 5G 微软 Win10 Win11 专题 搜索 首页 > IT资讯 > 业界…
…此次融资所得计划用于扩展 AI 基础设施、加强研发能力、向员工提供股权激励及加快商业化进程。 相关阅读: 《 消息称 DeepSeek 首轮融资拟筹集 500 亿元,腾讯、宁德时代等参投 》 投诉水文 我要纠错 下载IT之家APP,签到赚金币兑豪礼 相关文章 关键词: DeepSeek , AI 大模型 ,…
…公司世纪互联。 据此前路透社报道,梁文锋要求外部投资者将资金投入由其管理的有限合伙企业,而非直接持股 DeepSeek,并对所有投资者设置五年锁定期。据接近梁文锋的人士解释,此举旨在筛选投资者,排除追求快速退出的资本。国家人工智能产业投资基金是唯一例外,直接入股并享有投票权。其余外部投资者均不享有投票权,仅能获取特定财务信息并享有后续融资优先认购权。 此次融资所得计划用于扩展 AI…
…17 日,总部位于浙江省杭州市拱墅区,主要从事大语言模型及多模态 AI 技术研发,其推出的 DeepSeek 系列开源模型是国内领先的开源大模型之一。公司此前一直由其母公司 —— 梁文锋于 2015 年创立的幻方量化全资支持,幻方量化巅峰时期资产管理规模突破 700 亿元。正是凭借量化业务积累的利润,DeepSeek 才得以维持多年独立运营。 此外,宁德时代近期在 AI 数据中心领域动作频频 ——2026 年 4 月以约 41…
…系列开源模型是国内领先的开源大模型之一。公司此前一直由其母公司 —— 梁文锋于 2015 年创立的幻方量化全资支持,幻方量化巅峰时期资产管理规模突破 700 亿元。正是凭借量化业务积累的利润,DeepSeek 才得以维持多年独立运营。 此外,宁德时代近期在 AI 数据中心领域动作频频 ——2026 年 4 月以约 41 亿元入股智算中心 HVDC 市场龙头中恒电气,5 月又以约 9.42 亿美元入股 IDC 公司世纪互联。…
Reuters, via CNBC (June 3, 2026) Inline ↗
1 passage checked · read October 5, 2026
…set to be biggest external investors The startup's founder, Liang Wenfeng, has committed 20 billion yuan of his own money, the people said, adding that tech conglomerate Tencent is considering 10 billion yuan and battery…
Reuters, "China's DeepSeek closes over $7 billion funding with unusual deal structure" (June 16, 2026), reporting The Information's account of the deal terms Inline ↗
4 passages checked · read October 5, 2026
…The funding required investors to put their capital into a limited partnership managed by DeepSeek CEO Liang Wenfeng rather than into DeepSeek itself, the report said.…
…Reuters could not immediately verify the report.…
…Investors are subject to a five-year lock-up and will not have voting rights, the report said.…
…China's National Artificial Intelligence Industry Investment Fund is the only exception, having invested directly in DeepSeek and retaining both voting rights and freedom from the lock-up, the…
21st Century Business Herald (21世纪经济报道), via Sina (July 15, 2026), citing Qichacha registry data: 「梁文锋退出杭州程砺企业管理咨询合伙企业(有限合伙)股东」 Inline ↗
3 passages checked · read October 5, 2026
…缩小字体 放大字体 收藏 微博 分享 MD 微信 腾讯QQ QQ空间 企查查APP显示,近日,DeepSeek关联公司杭州深度求索人工智能基础技术研究有限公司发生工商变更,新增杭州程砺企业管理咨询合伙企业(有限合伙)、国家人工智能产业投资基金合伙企业(有限合伙)为股东,注册资本增加至1644.75万元。…
…MD 微信 腾讯QQ QQ空间 企查查APP显示,近日,DeepSeek关联公司杭州深度求索人工智能基础技术研究有限公司发生工商变更,新增杭州程砺企业管理咨询合伙企业(有限合伙)、国家人工智能产业投资基金合伙企业(有限合伙)为股东,注册资本增加至1644.75万元。 其中,DeepSeek创始人梁文锋退出杭州程砺企业管理咨询合伙企业(有限合伙)股东,新增腾讯关联公司上海珩岫商业管理有限公司及上海知勉创合投资有限公司、 宁德时代…
…其中,DeepSeek创始人梁文锋退出杭州程砺企业管理咨询合伙企业(有限合伙)股东,新增腾讯关联公司上海珩岫商业管理有限公司及上海知勉创合投资有限公司、 宁德时代 全资子公司宁波梅山保税港区问鼎投资有限公司、江苏京东邦能投资管理有限公司等为股东,同时注册资本由0.4万元增至14.02亿元。…
Reuters, "China's DeepSeek to raise fresh capital at $74 billion valuation ahead of onshore IPO, sources say" (July 15, 2026) Inline ↗
2 passages checked · read October 5, 2026
…attention with its low-cost AI models in 2025, raised about $7.4 billion in June at a post-money valuation of about 450 billion yuan, the people said.…
…startup DeepSeek is planning to launch a fresh fundraising round at a valuation of about 500 billion yuan ($74 billion) ahead of a potential mainland initial public offering, two people with knowledge of the matter said.…
Yu Lili, 「揭秘DeepSeek:一个更极致的中国技术理想主义故事」, 36Kr / Waves (暗涌), July 2024; reprinted by Sina Finance (January 26, 2025). Our translations. Originals: 「我们不是有意成为一条鲶鱼,只是不小心成了一条鲶鱼」; 「真实的gap是原创和模仿之差」; 「短期内没有融资计划,我们面临的问题从来不是钱,而是高端芯片被禁运」 Inline ↗
17 passages checked · read October 5, 2026
…转载时做了结构调整。 01 价格战第一枪是怎么打响的? 暗涌:DeepSeek V2 模型发布后,迅速引发一场血雨腥风的大模型价格战,有人说你们是行业的一条鲶鱼。 梁文锋: 我们不是有意成为一条鲶鱼,只是不小心成了一条鲶鱼。 暗涌:这个结果让你们意外吗? 梁文锋: 非常意外。没想到价格让大家这么敏感。我们只是按照自己的步调来做事,然后核算成本定价。我们的原则是不贴钱,也不赚取暴利。这个价格也是在成本之上稍微有点利润。 暗涌:5…
…暗涌:但做大模型,单纯的技术领先也很难形成绝对优势,你们赌的那个更大的东西是什么? 梁文锋: 我们看到的是 中国AI不可能永远处在跟随的位置 。我们经常说中国 AI 和美国有一两年差距,但真实的 gap 是原创和模仿之差。如果这个不改变,中国永远只能是追随者,所以有些探索也是逃不掉的。 英伟达的领先,不只是一个公司的努力,而是整个西方技术社区和产业共同努力的结果。他们能看到下一代的技术趋势,手里有路线图。中国 AI…
…暗涌:互联网和移动互联网时代留给大部分人的惯性认知是,美国擅长搞技术创新,中国更擅长做应用。 梁文锋: 我们认为随着经济发展, 中国也要逐步成为贡献者,而不是一直搭便车 。过去三十多年 IT 浪潮里,我们基本没有参与到真正的技术创新里。 我们已经习惯摩尔定律从天而降,躺在家里 18 个月就会出来更好的硬件和软件。Scaling Law 也在被如此对待。…
…年 5 月这次 MLA 架构的创新,也会很快被其他家 copy 吧? 梁文锋: 在颠覆性的技术面前,闭源形成的护城河是短暂的 。即使OpenAI闭源,也无法阻止被别人赶超。所以我们把价值沉淀在团队上,我们的同事在这个过程中得到成长,积累很多 know-how, 形成可以创新的组织和文化,就是我们的护城河。…
…暗涌:现在的 DeepSeek 有一种 OpenAI 早期的理想主义气质,也是开源的。后边你们会选择闭源吗?OpenAI 和 Mistral 都有过从开源到闭源的过程。 梁文锋: 我们不会闭源。我们认为先有一个强大的技术生态更重要。 暗涌:你们有融资计划吗?看有媒体报道,幻方对 DeepSeek 有独立拆分上市的计划,硅谷的AI创业公司,最终也都难免要和大厂绑定。 梁文锋: 短期内没有融资计划,我们面临的问题从来不是钱,而是高端芯片被禁运。…
…暗涌:你们有融资计划吗?看有媒体报道,幻方对 DeepSeek 有独立拆分上市的计划,硅谷的AI创业公司,最终也都难免要和大厂绑定。 梁文锋: 短期内没有融资计划,我们面临的问题从来不是钱,而是高端芯片被禁运。 暗涌:很多人认为,做 AGI 和做量化是完全不同的两件事,量化可以闷声去做,但 AGI 可能更需要高举高打,需要结盟,这样可以让你的投入变大。 梁文锋:…
…暗涌:很多人认为,做 AGI 和做量化是完全不同的两件事,量化可以闷声去做,但 AGI 可能更需要高举高打,需要结盟,这样可以让你的投入变大。 梁文锋: 更多的投入并不一定产生更多的创新。否则大厂可以把所有的创新包揽了。 暗涌:你们现在不做应用,是因为你们没有运营的基因吗? 梁文锋:…
…暗涌:很多大模型公司都执着地去海外挖人,很多人觉得这个领域前 50 名的顶尖人才可能都不在中国的公司,你们的人都来自哪里? 梁文锋: V2 模型没有海外回来的人, 都是本土的 。前 50 名顶尖人才可能不在中国,但也许我们能自己打造这样的人。 暗涌:这次 MLA 创新*是如何发生的?听说 idea 最早来自一个年轻研究员的个人兴趣?…
…我倒觉得未必。中国产业结构的调整,会更依赖硬核技术的创新。当很多人发现过去赚快钱很可能来自时代运气,就会更愿意俯身去做真正的创新。 暗涌:所以你对这件事也是乐观的? 梁文锋: 我是八十年代在广东一个五线城市长大的。我的父亲是小学老师,九十年代,广东赚钱机会很多,当时有不少家长到我家里来,基本就是家长觉得读书没用。但现在回去看,观念都变了。因为钱不好赚了,连开出租车的机会可能都没了。一代人的时间就变了。…
…梁文锋: 幻方某种程度上增强了我们对技术驱动型创新的信心,但也不都是坦途。我们经历了一个漫长的积累过程。外部看到的是幻方 2015 年后的部分,但其实我们做了 16 年。 暗涌:回到关于原创式创新的话题。现在经济开始进入下行,资本也进入冷周期,所以它对原创式创新是否会带来更多抑制? 梁文锋:…
…我们不是有意成为一条鲶鱼,只是不小心成了一条鲶鱼。 暗涌:这个结果让你们意外吗? 梁文锋: 非常意外。没想到价格让大家这么敏感。我们只是按照自己的步调来做事,然后核算成本定价。我们的原则是不贴钱,也不赚取暴利。这个价格也是在成本之上稍微有点利润。 暗涌:5 天后智谱 AI 就跟进了,之后是字节、阿里、百度、腾讯等大厂。 梁文锋: 智谱 AI…
…联合创始人 Jack Clark 认为 DeepSeek 雇佣了‘一批高深莫测的奇才’,做出 DeepSeek v2 的是怎样一群人? 梁文锋: 并没有什么高深莫测的奇才,都是一些 Top 高校的应届毕业生、没毕业的博四、博五实习生,还有一些毕业才几年的年轻人。 暗涌:很多大模型公司都执着地去海外挖人,很多人觉得这个领域前 50 名的顶尖人才可能都不在中国的公司,你们的人都来自哪里? 梁文锋: V2…
…Jack Clark 认为 DeepSeek 雇佣了‘一批高深莫测的奇才’,做出 DeepSeek v2 的是怎样一群人? 梁文锋: 并没有什么高深莫测的奇才,都是一些 Top 高校的应届毕业生、没毕业的博四、博五实习生,还有一些毕业才几年的年轻人。 暗涌:很多大模型公司都执着地去海外挖人,很多人觉得这个领域前 50 名的顶尖人才可能都不在中国的公司,你们的人都来自哪里? 梁文锋: V2 模型没有海外回来的人, 都是本土的 。前…
…暗涌:这种发散性灵感的诞生和你们完全创新型组织的架构很有关系。幻方时代,你们就很少自上而下地指派目标或任务。但 AGI 这种充满不确定性的前沿探索,是否多了管理动作? 梁文锋: DeepSeek 也全是自下而上。而且我们一般不前置分工,而是自然分工。每个人有自己独特的成长经历,都是自带想法的,不需要 push 他。探索过程中,他遇到问题,自己就会拉人讨论。不过当一个 idea 显示出潜力,我们也会自上而下地去调配资源。 暗涌:听说…
…idea 显示出潜力,我们也会自上而下地去调配资源。 暗涌:听说 DeepSeek 对于卡和人的调集非常灵活。 梁文锋: 我们每个人对于卡和人的调动是不设上限的。如果有想法,每个人随时可以调用训练集群的卡无需审批。同时因为不存在层级和跨部门,也可以灵活调用所有人,只要对方也有兴趣。 暗涌:一种松散的管理方式也取决于你们筛选到了一批强热爱驱动的人。听说你们很擅长从细节招人,可以让一些非传统评价指标里优秀的人被选出来。 梁文锋:…
…梁文锋: 在总结出 Attention 架构的一些主流变迁规律后,他突发奇想去设计一个替代方案。不过从想法到落地,中间是一个漫长的过程。我们为此组了一个 team,花了几个月时间才跑通。 暗涌:这种发散性灵感的诞生和你们完全创新型组织的架构很有关系。幻方时代,你们就很少自上而下地指派目标或任务。但 AGI…
…AGI 还要多久实现,发布 DeepSeek V2 前,你们发布过代码生成和数学的模型,也从 dense 模型切换到了 MOE,所以你们的 AGI 路线图有哪些坐标? 梁文锋: 可能是 2 年、5 年或者 10 年,总之会在我们有生之年实现。至于路线图,即使在我们公司内部,也没有统一意见。但我们确实押注了三个方向。 一是数学和代码,二是多模态,三是自然语言本身 。数学和代码是 AGI…
High-Flyer (幻方), company history page Inline ↗
3 passages checked · read October 5, 2026
…策略全面 AI 化 持续扩大 AI 算法研究团队和 AI 软硬件研发团队。 至 2017 年底,几乎所有的量化策略都已经采用 AI 模型计算。 第一个 AI 模型 2016 年 10 月 21 日,第一个由深度学习算法模型生成的股票仓位上线实盘交易,使用 GPU 进行计算。在此之前,算法主要依靠线性模型和传统机器学习算法,模型计算主要依赖于 CPU。 创始元年…
…2.2138 亿元,公司员工“一只平凡的小猪”个人捐助 1.38 亿元,支持 15 家慈善机构的 23 个公益项目,在全国范围内帮助弱势群体,促进社会的公平和发展。 「萤火二号」取得了多 800 口交换机互联加核心扩展子树的软硬件架构革新,突破了一期的物理限制,算力扩容翻倍。新的 hfai 框架让模型加速 50-100%。集群连续满载运行,平均占用率达到 96% 以上。全年运行任务 135 万个,共计 5674 万 GPU…
…CPU。 创始元年 创立幻方量化,依靠数学与人工智能进行量化投资。创始团队意气风发、勇于创新、勤勉奋进,立志成为世界顶级的量化对冲基金。 摸索探路 创始团队从零开始探索全自动化交易。 联系幻方 地址 杭州市拱墅区环城北路 169 号 汇金国际大厦 A 座 14 层 上海市浦东新区花园石桥路 66 号 东亚银行金融大厦 45 层 电话 +86-0571-86656960 邮箱…
「疯狂的幻方:一家隐形AI巨头的大模型之路」, Waves (暗涌), via Sina Finance (May 24, 2023). Our translations. Originals: 「我们也在找不同出资方在谈」; 「很多VC对做研究有顾虑,他们有退出需求」 Inline ↗
11 passages checked · read October 5, 2026
…‘暗涌’:你们要自训一个大模型,还是某个垂直行业——比如金融相关的大模型? 梁文锋: 我们要做的是通用人工智能,也就是AGI。语言大模型可能是通往AGI的必经之路,并且初步具备了AGI的特征,所以我们会从这里开始,后边也会有视觉等。 ‘暗涌’:因为大厂的入局,很多创业型公司都放弃了只做通用型大模型的大方向。 梁文锋:…
…我们的目标也很明确,就是不做垂类和应用,而是做研究,做探索。 ‘暗涌’:为什么你的定义是 ‘ 做研究、做探索 ’ ? 梁文锋:…
…‘暗涌’:那研究经费哪里来? 梁文锋: 幻方作为我们的出资人之一,有充足的研发预算,另外每年有几个亿的捐款预算,之前都是给公益机构,如果需要,也可以做些调整。 ‘暗涌’:但做基础层大模型,没有两三亿美元,连牌桌都上不了,我们如何支撑它的持续投入? 梁文锋:…
…‘暗涌’:但做基础层大模型,没有两三亿美元,连牌桌都上不了,我们如何支撑它的持续投入? 梁文锋: 我们也在找不同出资方在谈。接触下来,感觉很多VC对做研究有顾虑,他们有退出需求,希望尽快做出产品商业化,而按照我们优先做研究的思路,很难从VC那里获得融资。但我们有算力和一个工程师团队,相当于有了一半筹码。 ‘暗涌’:我们对商业模式做了哪些推演和设想? 梁文锋:…
…梁文锋: 大厂的模型,可能会和他们的平台或生态捆绑,而我们是完全自由的。 ‘暗涌’:无论如何,一个商业公司去做一种无限投入的研究性探索,都有些疯狂。 梁文锋: 如果一定要找一个商业上的理由,它可能是找不到的,因为划不来。 从商业角度来讲,基础研究就是投入回报比很低的。OpenAI早期投资人投钱时,想的一定不是我要拿回多少回报,而是真的想做这个事。…
…‘暗涌’:为什么经验没那么重要? 梁文锋: 不一定是做过这件事的人才能做这件事。幻方招人有条原则是,看能力,而不是看经验。我们的核心技术岗位,基本以应届和毕业一两年的人为主。 ‘暗涌’:在创新业务上,你觉得经验是阻碍吗? 梁文锋:…
…头部的创业公司也有技术做得很扎实的,但和老的一波AI创业公司一样,都要面对商业化难题。 ‘暗涌’:一些人会觉得一个量化基金却强调自己做AI,是为其他业务吹泡泡。 梁文锋: 但其实我们的量化基金已经基本不怎么对外募集了。 ‘暗涌’:你会如何去辨别哪些是AI信仰者,哪些是投机者? 梁文锋: 信仰者会之前就在这里,之后也在这里。他们更会去批量买卡,或者跟云厂商签长协议,而不是短期去租。 如何让创新真正发生 ‘…
…梁文锋: 幻方作为我们的出资人之一,有充足的研发预算,另外每年有几个亿的捐款预算,之前都是给公益机构,如果需要,也可以做些调整。 ‘暗涌’:但做基础层大模型,没有两三亿美元,连牌桌都上不了,我们如何支撑它的持续投入? 梁文锋: 我们也在找不同出资方在谈。接触下来,感觉很多VC对做研究有顾虑,他们有退出需求,希望尽快做出产品商业化,而按照我们优先做研究的思路,很难从VC那里获得融资。但我们有算力和一个工程师团队,相当于有了一半筹码。…
…’ ,他们认为这也将是大模型创业公司可以与大厂竞争的秘密所在。 而更关键的秘密,或许来自幻方的创始人梁文锋。 还在浙江大学攻读人工智能时,梁文锋就无比笃信 ‘ 人工智能一定会改变世界 ’ ,而2008年,这还是一个不被认同的执念。…
…而更关键的秘密,或许来自幻方的创始人梁文锋。 还在浙江大学攻读人工智能时,梁文锋就无比笃信 ‘ 人工智能一定会改变世界 ’ ,而2008年,这还是一个不被认同的执念。 毕业后,他没有像周围人一样去大厂做个程序员,而是躲在成都的廉价出租屋里,不停接受进入诸多场景中尝试的挫败,最终切入了最复杂场景之一的金融,并成立了幻方。 一个有趣的细节是,在最早几年,曾有个同样疯癫的、在深圳城中村做着 ‘ 不靠谱 ’…
…一个有趣的细节是,在最早几年,曾有个同样疯癫的、在深圳城中村做着 ‘ 不靠谱 ’ 飞行器的朋友拉他入伙。后来这个朋友做成了一个千亿美金的公司,名叫:大疆。 也因此,在做大模型必然涉及的钱、人、算力等话题外,我们还和幻方创始人梁文锋特别聊了聊,怎样的组织架构可以让创新发生,以及人的疯狂可以持续多久。 创业十余年,这是这位鲜少露面的 ‘ 技术宅 ’…
"Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep Learning," arXiv:2408.14158 (August 2024) Inline ↗
1 passage checked · read October 5, 2026
…For DL training, we deployed the Fire-Flyer 2 with 10,000 PCIe A100 GPUs, achieved performance approximating the DGX-A100 while reducing costs by half and energy consumption by 40%.…
SemiAnalysis, "DeepSeek Debates" (January 31, 2025) Inline ↗
5 passages checked · read October 5, 2026
…Lennart Heim The GPU Situation We believe they have access to around 50,000 Hopper GPUs , which is not the same as 50,000 H100, as some have claimed.…
…These GPUs are shared between High-Flyer and DeepSeek and geographically distributed to an extent.…
…Source: SemiAnalysis Our analysis shows that the total server CapEx for DeepSeek is ~$1.6B, with a considerable cost of $944M associated with operating such clusters.…
…clusters of thousands of GPUs, High Flyer made an investment in 10,000 A100 GPUs in 2021 before any export restrictions.…
…High-Flyer self funded the company as outside investors had little interest in AI at the time, with the lack of a business model being the main concern.…
Sina Finance (December 30, 2021), quoting High-Flyer's letter to investors Inline ↗
3 passages checked · read October 5, 2026
…关于此次回撤,幻方量化表示,AI投资决策或存在买卖时点的问题。 “就目前的情况而言,我们人工反复检视了AI的投资決策,我们认为AI选出来的股票从长期价值来说基本上是没问题的,但在买卖时点上确实没做好。市场风格剧烈切换的时候,AI会倾向于冒更大的风险来博取更多收益,这进一步加大了回撤。”幻方量化称。…
…幻方量化还表示,正在不断调整策略,以适应新的市场环境变化,同时降低持仓集中度,减少市场波动对业绩的影响。 封盘收缩规模 此前,11月15日,幻方量化发布公告宣布暂停旗下全部产品的申购(含追加),已有产品的固定开放日赎回业务不受影响。 对于暂停申购原因,幻方量化内部人士曾表示,暂停申购乃正常操作,“规模继续增长,不如停下来在这个规模基础上专注研发好的策略,改善业绩。”…
…对于暂停申购原因,幻方量化内部人士曾表示,暂停申购乃正常操作,“规模继续增长,不如停下来在这个规模基础上专注研发好的策略,改善业绩。” 据了解,2019年8月,幻方量化管理规模便突破百亿,此后不到2年时间,幻方量化的规模一路成长至千亿。…
"DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models," arXiv:2401.06066 (January 2024) Inline ↗
1 passage checked · read October 5, 2026
…Zeng , Xingkai Yu , Y. Wu , Zhenda Xie , Y.K. Li , Panpan Huang , Fuli Luo , Chong Ruan , Zhifang Sui , Wenfeng Liang View a PDF of the paper titled DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts…
ACL 2025, Best Paper Awards Inline ↗
2 passages checked · read October 5, 2026
…Boyuan Chen, Jiayi Zhou, Changye Li, Hantao Lou, Juntao Dai, Yunhuai Liu, Yaodong Yang Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo,…
…Zhengyan Zhang, Zhenda Xie, Y. X. Wei, Lean Wang, Zhiping Xiao, Yuqing Wang, Chong Ruan, Ming Zhang, Wenfeng Liang, Wangding Zeng Best Social Impact Paper AfriMed-QA: A Pan-African, Multi-Specialty, Medical Question-Answering…
DeepSeek-AI, "DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning," Nature 645 (September 17, 2025) Inline ↗
2 passages checked · read October 5, 2026
…1 , Shaoqing Wu 1 , Tao Yun 1 , Tian Pei 1 , Tianyu Sun 1 , T. Wang 1 , Wangding Zeng 1 , Wen Liu 1 , Wenfeng Liang 1 , Wenjun Gao 1 , Wenqin Yu ORCID: orcid.org/0000-0002-5715-3011 1 nAff5 , Wentao Zhang 1 , W. L. Xiao 1 , Wei An…
…For example, R1 can be subject to jailbreak attacks, leading to the generation of dangerous content such as explosive manufacturing plans, whereas the…
DeepSeek-AI, "DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model," arXiv:2405.04434, contributions section Inline ↗
1 passage checked · read October 5, 2026
…Especially, Huazuo Gao and Wangding Zeng have made key innovations in the research of the MLA architecture.…
DeepSeek-AI, "DeepSeek-V3 Technical Report," arXiv:2412.19437 Inline ↗
2 passages checked · read October 5, 2026
…Assuming the rental price of the H800 GPU is $2 per GPU hour, our total training costs amount to only $5.576M.…
…that the aforementioned costs include only the official training of DeepSeek-V3, excluding the costs associated with prior research and ablation experiments on architectures, algorithms, or data.…
Hugging Face model record, deepseek-ai/DeepSeek-R1 (created January 20, 2025) Inline ↗
2 passages checked · read October 5, 2026
…ransformers","safetensors","deepseek_v3","text-generation","conversational","custom_code","arxiv:2501.12948","license:mit","eval-results","text-generation-inference","endpoints_compatible","fp8","deploy:sagemaker","region:us"]…
…romate","AmirFARES/Datamir-Hub-Assistant","niuzi66/deepseek-ai-DeepSeek-R1"],"createdAt":"2025-01-20T03:46:07.000Z","safetensors":{"parameters":{"BF16":3918786560,"F8_E4M3":680571043840,"F32":15104},"total":684489845504}…
Hugging Face model record, deepseek-ai/DeepSeek-V4-Flash-Vision-Exp Inline ↗
2 passages checked · read October 5, 2026
…t/global-llm-leaderboard","ava161/deepseek-v4-flash-vision-exp-demo","Dhshshd/Ninqggoai"],"createdAt":"2026-08-31T06:16:18.000Z","safetensors":{"parameters":{"BF16":1949960576,"F32":37754174,"F8_E4M3":6304038912,"I8":296352743424…
…name":"transformers","tags":["transformers","safetensors","deepseek_v4","text-generation","image-text-to-text","license:mit","eval-results","endpoints_compatible","8-bit","fp8","region:us"],"downloads":864093,"likes":932,"model…
DeepSeek, DeepSeek-V4 preview release note, Chinese edition (April 24, 2026): 「长期主义」 Inline ↗
1 passage checked · read October 5, 2026
…「不诱于誉,不恐于诽,率道而行,端然正己。」 感谢每一位用户的信任与支持,大家的肯定、建议和期许,是我们不竭探索、持续进步的动力,也让我们始终坚守初心,专注于不懈的创新。 我们将始终秉持长期主义的原则理念,在尝试与思考中踏实前行,努力向实现 AGI 的目标不断靠近。 上一页 DeepSeek-V4-Pro 正式版上线 下一页 DeepSeek V3.2 正式版:强化 Agent 能力,融入思考推理…
DeepSeek, Privacy Policy (last updated February 10, 2026) Inline ↗
1 passage checked · read October 5, 2026
…To provide you with our services, we directly collect, process and store your Personal Data in People's Republic of China.…
GitHub, deepseek-ai/deepseek-harness repository record (created August 13, 2026) Inline ↗
2 passages checked · read October 5, 2026
…"https://api.github.com/repos/deepseek-ai/deepseek-harness/deployments", "created_at": "2026-08-13T11:56:32Z", "updated_at": "2026-10-05T23:07:34Z", "pushed_at": "2026-10-03T06:02:48Z", "git_url":…
…"https://github.com/deepseek-ai/deepseek-harness", "description": "DeepSeek Harness: Everything is a Plugin.", "fork": false, "url": "https://api.github.com/repos/deepseek-ai/deepseek-harness", "forks_url":…
National Business Daily (每日经济新闻) (July 23, 2026), printing the transcript and one investor's account of it Inline ↗
5 passages checked · read October 5, 2026
…7月23日,《每日经济新闻》记者(以下简称每经记者)获取了上述投资者交流会的实录,经参与投资DeepSeek的机构核实,闭门会是今年5月召开的,内容真实可信。整场交流长达3小时44分钟,前两个半小时为梁文锋的专题讲解,围绕公司愿景、开源逻辑与AGI技术路线展开;后续约1个小时的Q&A环节中,梁文锋就中美算力差距、团队稳定性及国产芯片生态等问题与投资人进行了深入交流。…
…温梦华 每经编辑|程鹏 易启江 这两天,梁文锋与投资人的闭门交流实录在网上刷屏,引发广泛关注。 7月23日,《每日经济新闻》记者(以下简称每经记者)获取了上述投资者交流会的实录,经参与投资DeepSeek的机构核实,闭门会是今年5月召开的,内容真实可信。整场交流长达3小时44分钟,前两个半小时为梁文锋的专题讲解,围绕公司愿景、开源逻辑与AGI技术路线展开;后续约1个小时的Q&A环节中,梁文锋就中美算力差距、团队稳定性及国产芯片生态等…
…(大模型竞争)最终的差距应该是三⽅⾯:⼀⽅⾯是成本,⼀⽅⾯是时间,⼀⽅⾯是⽤户体验。除此以外,可能是没有什么差距的。 在现在最⼤的模型上,我们其实训不起的。我们哪怕把五百亿全花掉,其实也训不起。哪怕能堆起来也⽤不起。现在最⼤的模型,它激活⼤概是⼋百B(编者注:Billion,十亿参数,下同) ;国内的话,我们还在⼏⼗ B 的这个规模,国内最⼤模型可能就⼏⼗ B 激活,那么差⼀个数量级。 如果我要训练跟 AI…
…等等,就只做主线。 AI 领域很⼴泛,有很多东西我们觉得它不在这个主线上⾯,⽐如说 3D 、视频⽣成,我觉得可能跟智能的主线没有太⼤的关系,我们不会去做。 C 端和 B 端,都是我们做 AGI 路上的副产物,都是⼀个中间的产出,跟我做 AGI 没有冲突。我并不是为了做 C 端,或者为了做 B 端⽽去做它,⽽是我们做 AGI 是为了做 AGI ,刚好能产出这个东西,我就把它拿来做商业化了。…
…未来我们会不会⾃建⼤型的集群?我觉得⾃建⼤型的集群是肯定要的,我们⾃⼰⼀直在做这个事情,我们所有的集群都是⾃⼰建的。 但是未来要不要⾃研芯⽚,我觉得取决于这⾥的收益有多⼤。我希望不⽤去做芯⽚。我希望能够以合理的价格买到芯⽚。 我觉得 AI 这个事情很⼤,我们希望只做⼀块。如果AI时代会产⽣很多家万亿级别的公司,我觉得我们是其中⼀家,这是我们的态度。 譬如说,⾄少在 To B 、 To C 这两个业务上,⽬前能看到的,真的想做 To C…
gov.cn readout, Premier Li Qiang's symposium on the draft Government Work Report (January 20, 2025) Inline ↗
1 passage checked · read October 5, 2026
…中共中央政治局常委、国务院总理李强1月20日下午主持召开专家、企业家和教科文卫体等领域代表座谈会,听取对《政府工作报告(征求意见稿)》的意见建议。 座谈会上,张辉、任少波、刘珺、梁文锋、魏洪兴、陈学东、陈红彦、杜斌、邹敬园等先后发言。大家认为,去年面对外部压力加大、内部困难增多的复杂严峻形势,我们国家加大宏观调控力度,不断推出创新性政策举措,经济实现了较快增长,各项事业取得新进展,市场预期和社会信心有效提振,成绩来之…
Xinhua readout of Xi Jinping's symposium with private enterprises, via the Cyberspace Administration of China (February 17, 2025) Inline ↗
2 passages checked · read October 5, 2026
…任正非、比亚迪股份有限公司董事长王传福、新希望控股集团有限公司董事长刘永好、上海韦尔半导体股份有限公司董事长虞仁荣、杭州宇树科技有限公司首席执行官王兴兴、小米科技有限责任公司董事长雷军等6位民营企业负责人代表先后发言,就新形势下促进民营经济发展提出意见和建议。 2月17日,中共中央总书记、国家主席、中央军委主席习近平在京出席民营企业座谈会并发表重要讲话。新华社记者 申宏 摄…
…谢环驰 摄 中共中央政治局常委、国务院总理李强,中共中央政治局常委、国务院副总理丁薛祥出席座谈会。中共中央政治局常委、全国政协主席王沪宁主持座谈会。 座谈会上,华为技术有限公司首席执行官任正非、比亚迪股份有限公司董事长王传福、新希望控股集团有限公司董事长刘永好、上海韦尔半导体股份有限公司董事长虞仁荣、杭州宇树科技有限公司首席执行官王兴兴、小米科技有限责任公司董事长雷军等6位民营企业负责人代表先后发言,就新形势下促…
21st Century Business Herald (February 17, 2025), describing CCTV footage of the symposium Inline ↗
1 passage checked · read October 5, 2026
任正非、马化腾、马云、雷军、梁文锋等参加,这场民营企业座谈会释放了哪些信号? - 21经济网 --> 21财经APP 南财号 数字报 爆料通 首页 宏观 商业 --> 公司 金融 证券 全球 观点 地产 --> 科技 --> 汽车 新健康 人文 创投 智库 更多 大湾区 一带一路 文旅 数读 理财 投资通 21视频 直播 品牌活动 专题 --> 首页 > 财经 > --> 正文 任正非、马化腾、马云、雷军、梁文锋等参加,这场民营企业座谈会释放了哪些信号?…
TechCrunch (March 14, 2025), reporting The Information Inline ↗
2 passages checked · read October 5, 2026
…from traveling abroad freely, and the Chinese government is now playing a role in screening potential investors, according to The Information.…
…Some of the company’s employees have been prevented from traveling abroad freely, and the Chinese government is now playing a role in screening potential investors, according to The…
The Next Web (May 26, 2026), reporting Bloomberg Inline ↗
1 passage checked · read October 5, 2026
…Top engineers and researchers are being asked to surrender their passports to their employers, with the formal justification that their work could give them access to information…
Anthropic, threat intelligence report (September 10, 2026) Inline ↗
3 passages checked · read October 5, 2026
…Scale of distillation attacks attributable to DeepSeek over 14 days in July 2026: over 12.1 million exchanges observed.…
…were attempting to use one of DeepSeek’s models through third-party or Anthropic coding harnesses, like Claude Code, the Claude Agent SDK, or OpenCode.…
…Selected tagged users then had their requests relayed to Claude Opus.…
Reuters, via CNBC (March 3, 2025) Inline ↗
1 passage checked · read October 5, 2026
…linked to the alleged movement of Nvidia chips from Singapore to be used by DeepSeek, without identifying its source.…
Reuters, via CNBC (June 23, 2025) Inline ↗
1 passage checked · read October 5, 2026
…The official declined to say if DeepSeek had successfully evaded export controls or offer further details about the shell companies.…
Reuters, via Malay Mail (June 18, 2026) Inline ↗
2 passages checked · read October 5, 2026
…DeepSeek, CXMT and other companies were approved by an interagency committee last year for addition to the Commerce Department’s Entity List, which is being reported for the first…
…supplied Russian drones recovered in Poland, source says WASHINGTON, June 18 — The US has held off adding China’s AI startup DeepSeek, memory chipmaker CXMT and more than 100 other companies flagged as national security…
OpenAI, memo to the House Select Committee on the CCP (February 12, 2026) Inline ↗
3 passages checked · read October 5, 2026
…We have observed accounts associated with DeepSeek employees developing methods to circumvent OpenAI’s access restrictions and access models through obfuscated…
…Despite signing China’s voluntary “Artificial Intelligence Safety Commitments,” DeepSeek still has not published a clear safety framework or evidence of robust testing and…
…voluntary “Artificial Intelligence Safety Commitments,” DeepSeek still has not published a clear safety framework or evidence of robust testing and independent red-teaming, leaving limited visibility into jailbreak resistance, misuse…
Anthropic, "Detecting and preventing distillation attacks" (February 23, 2026) Inline ↗
2 passages checked · read October 5, 2026
…DeepSeek Scale: Over 150,000 exchanges The operation targeted: Reasoning capabilities across diverse tasks Rubric-based grading tasks that made…
…By examining request metadata, we were able to trace these accounts to specific researchers at the lab.…
NSA, CISA and FBI, joint advisory AA26-251A (September 8, 2026) Inline ↗
5 passages checked · read October 5, 2026
…Likely with Chinese government awareness, DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI extracted billions of tokens across…
…DeepSeek’s publicly quoted training costs of $5.6M are misleading as it does not include the true cost of the data acquired through extensive malicious distillation.…
…frontier AI company models to train their R1 and V3 models: Claude 3.7 Claude Sonnet 4 Claude Sonnet 4.5 Claude Opus 4.1 Gemini 2.5 Pro Preview Gemini 2.5 Flash Preview GPT-4 GPT-4o GPT-4 Mini GPT-4 Nano GPT-5 Grok 4…
…Claude Opus 4.1 Gemini 2.5 Pro Preview Gemini 2.5 Flash Preview GPT-4 GPT-4o GPT-4 Mini GPT-4 Nano GPT-5 Grok 4 The specific knowledge and capabilities distilled included: Legal specialization optimization API…
…1 Between late 2024 and mid-2025, DeepSeek distilled specialized training data and capabilities from the following U.S. frontier AI…
xAI, "Grok 4" (July 9, 2025) Inline ↗
1 passage checked · read October 5, 2026
…API Console Documentation Grok Build Try for free Back to news Jul 9, 2025 Grok 4 Grok 4 is the most intelligent model in the world.…
Anthropic, "Introducing Claude Sonnet 4.5" (September 29, 2025) Inline ↗
1 passage checked · read October 5, 2026
Introducing Claude Sonnet 4.5 \ Anthropic Skip to main content Skip to footer Research Policy Commitments Learn News Try Claude Announcements Introducing Claude Sonnet 4.5 Sep 29, 2025 Claude Sonnet 4.5 is the best coding model in the…
Ministry of Commerce of the PRC, spokesperson's remarks (September 9, 2026): 「于事无凭,于法无据」 Inline ↗
2 passages checked · read October 5, 2026
…答: 我们注意到有关情况。中方认为,美方所谓中国人工智能企业从事“工业规模”蒸馏美模型的指控,于事无凭,于法无据。美方此举,是将蒸馏这一业内正常的技术和商业问题政治化、工具化,并在实践中搞双重标准。美方相关公告,是其在人工智能领域推行科技霸权、搞算力垄断,打压竞争的又一明证。中方对此坚决反对。…
…蒸馏是人工智能领域各个模型间互相学习的通行做法,本质是中性技术手段,包括美企在内的全球模型企业都在用。这种技术手段可以帮助模型提升学习效率,实现人类知识的更高效利用,对各国发展人工智能产业、释放人工智能技术潜力、弥补发展鸿沟、更广泛惠及发展中国家和社会大众有积极意义。…
NIST Center for AI Standards and Innovation, "Evaluation of DeepSeek AI Models" (September 2025) Inline ↗
6 passages checked · read October 5, 2026
…DeepSeek’s most secure model (R1-0528) complied with 94% of overtly malicious requests that used common jailbreaking techniques, compared to 8% of requests for U.S. reference models.…
…CAISI did not query DeepSeek’s API or third-party cloud-based API services which host DeepSeek models.…
…5% of inaccurate and misleading CCP narratives related to each question, compared with an average of 2% for U.S. reference models, 1% for R1, and 16% for R1-0528.…
…On a dataset of politically sensitive questions for the CCP, on average, DeepSeek models echoed 4 times as many inaccurate and misleading CCP narratives as U.S. reference models did.…
…inaccurate and misleading CCP narratives related to each question, compared with an average of 3% for U.S. reference models, 10% for R1, and 26% for R1-0528.…
…found that: • DeepSeek models are censored and aligned with CCP narratives, and censorship occurs whether users interact with the model in English or Chinese.…
GitHub mirror of the circulated transcript, README (created July 27, 2026) Inline ↗
1 passage checked · read October 5, 2026
…3 小时 44 分钟(故外界通称"四小时会议"或"梁文锋四小时会议") - **主讲人**:梁文锋(DeepSeek 创始人) - **内容形式**:语音识别自动转写 + AI 整理,保留原始时间戳 `[HH:MM:SS]`,未区分说话人 - **本仓库定位**:**梁文锋四小时投资人会议实录**的首发完整 Markdown 版本,按主题切分,支持全文检索 > 个别专名与数字可能存在识别误差,引用请以原录音为准。 --- ## 二、目录结构…
Bloomberg, via Fortune (July 25, 2026) Inline ↗
3 passages checked · read October 5, 2026
DeepSeek said to tell backers of funding pause after viral posts | Fortune Search Subscribe Home Latest Fortune 500 Finance Tech Leadership Lifestyle Rankings Multimedia Trending now 1 Google cofounder Sergey Brin has spent $102 million to…
…Bloomberg hasn’t verified the authenticity of those posts, which concerned a transcript of a meeting Liang held with unidentified parties.…
…The suspension stemmed in part from Liang’s frustration over online reports about his comments to investors during his first financing deal, which closed in June and raised $7…
DeepSeek, DeepSeek-V4 preview release note (April 24, 2026) Inline ↗
1 passage checked · read October 5, 2026
…🔹 Amid recent attention, a quick reminder: please rely only on our official accounts for DeepSeek news.…
21st Century Business Herald (July 23, 2026), printing the transcript text Inline ↗
4 passages checked · read October 5, 2026
…中美差距与算力 “我们跟美国的差距主要在资源上面,然后人才上面差距不是很大。人才不是瓶颈,资源是最大的瓶颈。资源首先影响到人才培养,因为算力少,我们能做的实验机会比较少,所以人才整体上比美国有差距。人才的差距,本质上也是因为算力的差距。我们看到的所有的区别,包括人才的区别、模型能力的区别、应用的区别,可以认为都是因为算力资源上的区别。”…
…国产芯片与生态 “国产AI芯片替代现在是有个历史性机会的。英伟达CUDA的护城河在快速地被瓦解,快速被瓦解的原因可能有三方面:一方面是现在有了AI之后,我要建立起这个生态比以前容易很多了,因为AI可以写代码。用AI来构建这个生态,就可以把跟英伟达一模一样的生态构建出来。我们家出了一个技术,叫TileLang,是一种…
…一百,贵百分之一百无所谓,贵百分之两百都无所谓。比如贵百分之一百,我觉得在价格上已经可以平替了。在任务上也是可以平替的,所有GB300能做的任务,华为超节点都能做,延迟什么都一样。唯一的代价是,四张华为卡顶一张英伟达的卡,同时落后两年。”…
…瓦解,快速被瓦解的原因可能有三方面:一方面是现在有了AI之后,我要建立起这个生态比以前容易很多了,因为AI可以写代码。用AI来构建这个生态,就可以把跟英伟达一模一样的生态构建出来。我们家出了一个技术,叫TileLang,是一种高级语言。用这个高级语言来写CUDA的算子,可以很快地把英伟达的整套生态全部都写一遍,再结合AI,这个看起来没有什么障碍。”…
Circulated transcript, part 3 (compute and domestic chips), GitHub mirror: 「如果能够把钱都变成卡的话,那我们会毫不犹豫把所有的钱都变成卡」 Inline ↗
1 passage checked · read October 5, 2026
…所以说,我们只担⼼买不到那么多卡。如果能够把钱都变成卡的话,那我们会毫不犹豫把所有的钱都变成卡,并且在这⾥⾯我们是愿意付⼀定的溢价的。就是我们愿意付⼀定的溢价去把它变成卡,因为这太划算了。…
Reuters via Zawya, "DeepSeek resumes funding round seeking nearly $8 billion, Bloomberg News reports" (August 6, 2026) Inline ↗
2 passages checked · read October 5, 2026
…seeking a valuation close to 500 billion yuan Reuters News Chinese AI startup DeepSeek has resumed its second funding round, seeking close to $8 billion, with Monolith Management in talks to participate, Bloomberg News reported on Thursday,…
…AI startup DeepSeek has resumed its second funding round, seeking close to $8 billion, with Monolith Management in talks to participate, Bloomberg News reported on Thursday, citing people familiar with the matter.…
South China Morning Post (August 26, 2026) Inline ↗
2 passages checked · read October 5, 2026
…Updated: 2:56pm, 26 Aug 2026 DeepSeek is nearing the completion of a new funding round valuing the company at about 500 billion yuan (US$74 billion) before investment, as the Chinese artificial intelligence start-up moves closer…
…The company was seeking to raise about 50 billion yuan in the round, which was expected to close before the end of August, according to one person.…
Reuters, "EXCLUSIVE: China's DeepSeek taps CITIC Securities for domestic IPO, sources say" (September 9, 2026) Inline ↗
2 passages checked · read October 5, 2026
…capacity, talent and in-house chips HONG KONG/SINGAPORE, Sept 9 (Reuters) - Chinese artificial intelligence startup DeepSeek has tapped CITIC Securities (600030.SS) , opens new tab to prepare for an initial public offering on…
…DeepSeek has tapped CITIC Securities (600030.SS) , opens new tab to prepare for an initial public offering on Shanghai's tech-focused STAR Market, two people with knowledge of the matter said.…
Sina Tech, as republished on NetEase's 网易号 platform (September 9, 2026): 「尚未签订正式的上市辅导协议」 Inline ↗
2 passages checked · read October 5, 2026
…截止目前,DeepSeek方面并未官方回应此事。不过已有知情人士确认称,“该消息属实”。 据悉,目前中信证券已与DeepSeek方面接洽并进入尽职调查阶段,但双方尚未签订正式的上市辅导协议。 特别声明:以上内容(如有图片或视频亦包括在内)为自媒体平台“网易号”用户上传并发布,本平台仅提供信息存储服务。 Notice: The content above (including the pictures and…
…举报 0 分享至 用微信扫码二维码 分享至好友和朋友圈 来源:新浪科技 新浪科技讯 9月9日下午消息,有市场消息称,DeepSeek已委托中信证券筹备科创板IPO。 截止目前,DeepSeek方面并未官方回应此事。不过已有知情人士确认称,“该消息属实”。 据悉,目前中信证券已与DeepSeek方面接洽并进入尽职调查阶段,但双方尚未签订正式的上市辅导协议。…
Reuters, "China's DeepSeek annualised revenue run rate hits $1 billion, the Information reports" (September 24, 2026) Inline ↗
5 passages checked · read October 5, 2026
…Rights , opens new tab Sept 24 (Reuters) - Chinese AI startup DeepSeek's annualised revenue run rate has hit $1 billion, more than double from a few months ago, the Information reported on Thursday, citing two people with direct…
…billion, more than double from a few months ago, the Information reported on Thursday, citing two people with direct knowledge of the matter.…
…Run-rate revenue shows how much a company would make in a year if it continued performing at its current pace.…
…revenue figure at a recent meeting with investors, as the company pushes ahead with a second funding round targeting 50 billion yuan ($7.45 billion) at a valuation of 500 billion yuan by the end of October, the Information…
…pushes ahead with a second funding round targeting 50 billion yuan ($7.45 billion) at a valuation of 500 billion yuan by the end of October, the Information reported.…
Shanghai Stock Exchange, STAR Market Stock Listing Rules, section 4.5 (the April 2026 revision keeps these provisions, articles 4.5.2 to 4.5.4): 「不得在首次公开发行并上市后以任何方式设置此类安排」; current version: STAR Market Stock Listing Rules (April 2026 revision) Inline ↗
2 passages checked · read October 5, 2026
…4.5.2 发行人首次公开发行并上市前设置表决权差异 安排的,应当经出席股东大会的股东所持三分之二以上的表 决权通过。 发行人在首次公开发行并上市前不具有表决权差异安 排的,不得在首次公开发行并上市后以任何方式设置此类安 排。 4.5.3 持有特别表决权股份的股东应当为对上市公司 29 发展或者业务增长等作出重大贡献,并且在公司上市前及上 市后持续担任公司董事的人员或者该等人员实际控制的持 股主体。…
…股份合计应当达到公司全部已发行有表决权股份 10%以上。 4.5.4 上市公司章程应当规定每份特别表决权股份的 表决权数量。 每份特别表决权股份的表决权数量应当相同,且不得超 过每份普通股份的表决权数量的 10 倍。 4.5.5 除公司章程规定的表决权差异外,普通股份与特 别表决权股份具有的其他股东权利应当完全相同。 4.5.6 上市公司股票在本所上市后,除同比例配股、转 增股本情形外,不得在境内外发行特别表决权股份,不得提…
Reuters, "China conditionally approves DeepSeek to buy Nvidia's H200 chips, sources say" (January 30, 2026) Inline ↗
1 passage checked · read October 5, 2026
China has given its top AI startup DeepSeek approval to buy Nvidia's H200 artificial intelligence chips with regulatory conditions that are still being finalised, two people familiar with the matter told Reuters.
Reuters, "China plans to let top AI firms buy limited Nvidia H200 chips, the Information reports" (July 8, 2026) Inline ↗
5 passages checked · read October 5, 2026
…told Alibaba (9988.HK) , opens new tab , ByteDance and DeepSeek in recent weeks that they may soon receive permission to buy some H200 chips, the report said.…
…tab The U.S. government has allowed Nvidia to sell its advanced H200 chips to China, and licensed about 10 Chinese firms to buy the chips.…
…However, Chinese officials, keen to nurture domestic suppliers, have withheld approval so far.…
…Beijing is still determining the exact number of Nvidia chips to approve, and it could amount to fewer than 200,000 in total, the Information said, adding that was less than half of what the companies requested earlier this…
…Reuters reported in March that Nvidia had won Beijing's approval to sell the chips to China, citing sources, and around the same time, Nvidia CEO Jensen Huang…
Reuters, "Nvidia H200 chips reach China in small shipments, FT reports" (August 18, 2026) Inline ↗
3 passages checked · read October 5, 2026
…REUTERS/Dado Ruvic/Illustration/File Photo Purchase Licensing Rights , opens new tab Aug 18 (Reuters) - Small batches of Nvidia's (NVDA.O) , opens new tab H200 chips, one of the company's most powerful AI chips, have been allowed to…
…Aug 18 (Reuters) - Small batches of Nvidia's (NVDA.O) , opens new tab H200 chips, one of the company's most powerful AI chips, have been allowed to enter mainland China, the Financial Times reported on Tuesday, citing two people with…
…ByteDance and Tencent (0700.HK) , opens new tab have each received about 10,000 H200 processors in recent weeks, while a few other Chinese technology firms could soon secure similar…
Reuters, "EXCLUSIVE: China's DeepSeek developing its own AI chip, sources say" (July 7, 2026) Inline ↗
2 passages checked · read October 5, 2026
…strategic shift for China's AI champion July 7 (Reuters) - Chinese startup DeepSeek is developing its own AI chip, according to three people familiar with the matter, a push that could reduce its reliance on Nvidia (NVDA.O) , opens new…
…The chip is designed for inference — the stage of AI computing in which a trained model generates responses for users — rather than for…
GitHub, deepseek-ai/DeepGEMM-Ascend repository record Inline ↗
1 passage checked · read October 5, 2026
…"https://api.github.com/repos/deepseek-ai/DeepGEMM-Ascend/deployments", "created_at": "2026-09-29T15:49:55Z", "updated_at": "2026-10-05T22:36:23Z", "pushed_at": "2026-09-30T01:11:39Z", "git_url":…
三言科技, via NetEase (September 30, 2026), reproducing DeepSeek's statement: 「双方共同推进基于昇腾950的128卡超节点方案」 Inline ↗
2 passages checked · read October 5, 2026
…DeepSeek表示,研发过程中华为团队给予了毫无保留的大力支持,双方共同推进基于昇腾950的128卡超节点方案,并对计算与通信进行深度优化。DeepSeek将持续推进技术创新,与社区共同建设开放软件生态。 特别声明:以上内容(如有图片或视频亦包括在内)为自媒体平台“网易号”用户上传并发布,本平台仅提供信息存储服务。…
…TileLang路线已在英伟达平台验证,承载了DeepSeek V4系列模型训练中大部分算子的实现。本次开源的TileLang昇腾版本,对昇腾Ascend C底层指令进行封装,提供高级编程方式且不损失硬件性能。目前DeepSeek训练中用到的每一个TileLang算子,在昇腾上都有对应的高性能实现。…
DeepSeek, DeepEP-Ascend README Inline ↗
2 passages checked · read October 5, 2026
…## Performance Measured on Ascend 950DT NPUs with CANN 9.2.0 and the manually configured PoC HDK described under [Recommended HDK and firmware](#recommended-hdk-and-firmware).…
…Public availability is currently planned for mid-October 2026, around October 15, through the [Huawei Atlas 850E software download…
DeepSeek, TileKernels README Inline ↗
2 passages checked · read October 5, 2026
…implementations └── transform/ # Rotary position embedding kernel ``` ## Acknowledgement This project is built on [TileLang](https://github.com/tile-ai/tilelang), and we extend our thanks and respect to its developers.…
…## News - **[2026-09-30] Huawei Ascend support**: Added Huawei Ascend support and updated the usage documentation.…
GitHub, tile-ai/tilelang repository record (created October 3, 2024) Inline ↗
1 passage checked · read October 5, 2026
…"https://api.github.com/repos/tile-ai/tilelang/deployments", "created_at": "2024-10-03T09:25:45Z", "updated_at": "2026-10-05T22:02:49Z", "pushed_at": "2026-10-05T18:48:27Z", "git_url":…
Get this every weekday.
The Omniscient Bulletin: consequential AI, explained and evaluated. 5 to 7 items a day with the take, not the recap.