{"id":56046,"date":"2026-08-20T09:16:29","date_gmt":"2026-08-20T01:16:29","guid":{"rendered":"https:\/\/www.1ai.net\/?p=56046"},"modified":"2026-08-20T09:16:29","modified_gmt":"2026-08-20T01:16:29","slug":"%e6%99%ba%e8%b0%b1%e5%94%90%e6%9d%b0%ef%bc%9a%e6%a8%a1%e5%9e%8b%e6%80%bb%e5%8f%82%e6%95%b0%e8%be%be%e5%88%b0%e7%9f%a5%e8%af%86%e9%98%88%e5%80%bc%e5%90%8e%ef%bc%8c%e8%83%bd%e5%8a%9b%e8%b7%83%e5%8d%87","status":"publish","type":"post","link":"https:\/\/www.1ai.net\/en\/56046.html","title":{"rendered":"Tang Jie: Once the total parameters of the model reach the knowledge threshold, the ability leap should shift to Post-training"},"content":{"rendered":"<p>News on August 20,<a href=\"https:\/\/www.1ai.net\/en\/tag\/%e6%99%ba%e8%b0%b1\" title=\"[View articles tagged with [Smart Spectrum]]\" target=\"_blank\" >Zhipu<\/a> AI CHIEF SCIENTIST TANG JIE POST STATED THAT THE GLM-53 WOULD BE SCALED UP THROUGH EXPANDED TRAINING<a href=\"https:\/\/www.1ai.net\/en\/tag\/%e6%a8%a1%e5%9e%8b\" title=\"_Other Organiser\" target=\"_blank\" >Model<\/a>CAPABILITY, RATHER THAN SIMPLY INCREASING THE NUMBER OF PARAMETERS. THE MODEL USES THE SAME BASE MODEL, STRUCTURE, TOTAL PARAMETER AND ACTIVATED PARAMETER LEVELS AS GM-5.2, BUT AFTER ONE MONTH OF LONG-CYCLE ENVIRONMENTAL TRAINING AND INTENSIVE LEARNING, THE MODEL PERFORMANCE HAS IMPROVED SIGNIFICANTLY\u3002<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-56047\" title=\"941d1639j00tk1nie005ed000gg00mgmmmm\" src=\"https:\/\/www.1ai.net\/wp-content\/uploads\/2026\/08\/941d1639j00tk1nie005ed000gg00mgm.jpg\" alt=\"941d1639j00tk1nie005ed000gg00mgmmmm\" width=\"592\" height=\"808\" \/><\/p>\n<p>Tang Jie recalled in his post that early studies had concluded that models should expand the parameters at a faster pace, but that follow-up studies had shown that, in the case of limited computing resources, a more balanced ratio between training data and model parameters was needed. As models are used on a large scale, the proportion of reasoning costs in the total cost of the model ' s life cycle is increasing, and smaller-trained models with more data are a viable direction\u3002<\/p>\n<p>He also mentioned that the thin mix expert model needed to distinguish between total and activated parameters: While the former broadly determine the extent of knowledge that the model can accommodate, the latter and the depth of effective reasoning affect the model ' s ability to complete long-chain reasoning. In the case of software loopholes, this is not just a search of known loopholes, but a requirement for the model to continue to complete multistep reasoning\u3002<\/p>\n<p>IN TANG JIE ' S VIEW, WHEN THE TOTAL NUMBER OF MODEL PARAMETERS REACHES A THRESHOLD OF SUFFICIENT TO CARRY A WIDE RANGE OF KNOWLEDGE, FURTHER ENHANCEMENT OF CAPACITY CAN SHIFT TO EFFECTIVE REASONING DEPTH AND POST-TRAINING. THIS GM-5.3 IS A CONTROLLED EXPERIMENT TO THIS JUDGEMENT, BUT HE ALSO INDICATES THAT THERE IS ROOM FOR EXPANSION OF THE BASE MODEL SCALE, PRE-TRAINED DATA VOLUME AND SINGLE-WAY CALCULATIONS, AND THAT FUTURE WORK MAY CONTINUE IN THE DIRECTION OF ONGOING TRAINING AND PRE-TRAINING\u3002<\/p>","protected":false},"excerpt":{"rendered":"<p>ACCORDING TO THE 20 AUGUST MESSAGE, TANG JIE, THE CHIEF SCIENTIST OF THE BRAIN SPECTRA AI, GLM-5.3 WILL UPGRADE MODEL CAPABILITIES THROUGH EXPANDED TRAINING, RATHER THAN SIMPLY INCREASING THE NUMBER OF PARAMETERS. THE MODEL USES THE SAME BASE MODEL, STRUCTURE, TOTAL PARAMETER AND ACTIVATED PARAMETER LEVELS AS GM-5.2, BUT AFTER ONE MONTH OF LONG-CYCLE ENVIRONMENTAL TRAINING AND INTENSIVE LEARNING, THE MODEL PERFORMANCE HAS IMPROVED SIGNIFICANTLY. TANG JIE RECALLED IN HIS POST THAT EARLY STUDIES HAD CONCLUDED THAT MODELS SHOULD EXPAND THE PARAMETERS AT A FASTER PACE, BUT THAT FOLLOW-UP STUDIES HAD SHOWN THAT, IN THE CASE OF LIMITED COMPUTING RESOURCES, A MORE BALANCED RATIO BETWEEN TRAINING DATA AND MODEL PARAMETERS WAS NEEDED. AS THE MODEL IS USED ON A LARGE SCALE, THE PROPORTION OF REASONING COSTS IN THE TOTAL COST OF THE MODEL ' S LIFE CYCLE INCREASES, AND SMALLER-TRAINED MODELS WITH MORE DATA BECOME ONE<\/p>","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[146],"tags":[2680,1489],"collection":[],"class_list":["post-56046","post","type-post","status-publish","format-standard","hentry","category-news","tag-2680","tag-1489"],"acf":[],"_links":{"self":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/posts\/56046","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/comments?post=56046"}],"version-history":[{"count":0,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/posts\/56046\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/media?parent=56046"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/categories?post=56046"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/tags?post=56046"},{"taxonomy":"collection","embeddable":true,"href":"https:\/\/www.1ai.net\/en\/wp-json\/wp\/v2\/collection?post=56046"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}