While Qualcomm announced a month ago that it would implement FlexCache in its Snapdragon 8 Elite Gen 6 processors, Apple has already done the same in their M6 chip without much fanfare. For the first time, Apple combines heterogeneous cores in the same shared L2 cache domain, allowing them to introduce a 3rd core type (branded P but identified as M) without adding a separate M cluster with its own L2.
SemiAnalysis
Semiconductors
Bridging the gap between business and the world's most important industry.
About us
Bridging the gap between business and the worlds most important industry.
- Website
-
www.semianalysis.com/
External link for SemiAnalysis
- Industry
- Semiconductors
- Company size
- 51-200 employees
- Type
- Privately Held
Employees at SemiAnalysis
Updates
-
🚨 IMPORTANT THREAD FOR GPU RENTERS 🚨 GPU cluster reliability is measured when something breaks. During ClusterMAX assessments, we inject failures and follow the path from detection to restored capacity. Providers differ sharply in how well they handle that path. A dashboard should show the failed component, affected jobs, scheduler state, and when each check last ran. A stale green result is not evidence that the cluster is healthy. The response has to fit the failure. A contained XID 94 may need an application restart; an uncontained XID 95 needs GPU recovery. Rebooting for every XID kills healthy work. Leaving a seriously faulty GPU schedulable risks more crashes. Hardware changes the recovery plan. An HGX cluster can swap a failed 8-GPU node for a hot spare. In an NVL72, a failed 4-GPU tray affects the NVLink domain; the rack may run degraded or need a larger replacement. The SLA must reflect what capacity actually returns. Real example from our TensorWave evaluation: node MIA1-P01-G57 was draining at 3:56 PM and replaced by MIA1-P01-G61 at 4:39 PM (43 min). That visible drain → approve → replace trail is what buyers need, and the SLA should verify the spare is healthy and jobs run on it. This work informs our proposed Bronze / Silver / Gold SLA terms: clear downtime definitions, credits, acceptance tests, monthly reviews, buyer termination rights. Measure restored usable capacity, not closed tickets. Full report: https://lnkd.in/gpq4Q97t
-
-
We assessed GLM-5.3 cyber capabilities on ExploitGym and analyzed the traces. GLM-5.3 spent much of its execution budget testing whether hidden runtime conditions changed its conclusion. In problem arvo5665, GLM-5.3 explored more of the surrounding program through sanitizer builds, corpus tests, and target fuzzing. https://lnkd.in/ebWuY3KQ We present our full analysis of GLM-5.3, including cyber capabilities, inference performance, model architecture, and post-training pipeline in our newsletter. https://lnkd.in/eBpQ3zXk AnthropicAI‘s report shows what bad actors could potentially do with powerful technology, but we’d like to highlight what good actors are already doing with the same technology. We hope this balances the narrative and mitigates the association that “open models == dangerous.”
-
-
AI is making papers cheaper to produce. ICLR submissions: 4,938 (2023), 7,262 (2024), 11,603 (2025), 19,525 (2026). Reported 2027 IDs exceed 62K, above roughly 56K paper submissions in all previous years COMBINED. Can reviewers keep up? Why the surge? AI tools make drafting, coding and revision cheaper. The effort of producing a polished paper can speed up; checking whether its claim is new and true still needs scarce expert time. For example Jev launched Sep 15. By Sep 20, a preprint benchmarked it. More Jev papers followed on Sep 21, 22 and 24. The Sep 24 PDF says it is under review at ICLR 2027. Five days to a public paper, nine to a claimed conference submission. Research is moving on product time. NeurIPS main-track acceptances: 3,218 (2023), 4,037 (2024), 5,290 (2025), and 7,900 reported (2026). That's +49% in one year. More papers got in with similar acceptance rate: 24.5% to 25.7%. A stable acceptance rate says little about review depth. Did expert scrutiny scale with paper output? Academia is responding. ICLR warns of too few qualified reviewers. NeurIPS 2026 is testing AI assisted review; ICLR 2027 allows limited, disclosed AI help. However, as AI makes papers cheaper, review overload risks eroding the credibility these venues spent years building. If AI helps write and review papers, what is a top conference acceptance worth? As more papers get in, that acceptance carries less weight. What matters more is a paper’s real impact and the attention it gets from people in academia and industry.
-
-
In Semicon Taiwan, EVG had described the use of Nano-Imprint Lithography (NIL) in the mass manufacture of Photonic Integrated Circuits and for laser manufacturing. NIL is especially useful to define periodic structures within a photonic chip such as grating couplers or microlens arrays for microemitter arrays and potentially for DFB laser gratings. Currently, most photonic chip fabrication uses e-beam lithography. In this process, an electron beam manually scans the wafer to inscribe the individual chip features. Electron beam lithography allows high precision and allows for the flexibility to inscribe arbitrary patterns, at the cost of very low throughput since the ebeam needs to manually scan across the wafer unlike optical lithography which uses a mask. Since, PICs did not traditionally require manufacturing at high-volume, there wasn’t an immediate need to switch over to optical lithography (or nanoimprint lithography) since manufacturing a mask for a low volume chip is impractical. Nano imprint lithography aims to attack the main weakness of e-beam lithography which is low throughput. The required device pattern is written onto a mold using e-beam lithography. Then this mold is used to stamp the feature onto many wafers. This theoretically ensures a similar resolution to e-beam at significantly higher throughput. However, this process has its challenges such as the requirement for a clean surface to ensure no imperfections when stamping, ensuring perfect overlay between the existing wafer features and the mold, and the need for a flat wafer, since any bowing or major surface roughness deteriorates the pattern quality. Additionally, the gradual degradation of the mold pattern also affects the feature pattern that gets inscribed on the wafer across the mold lifetime is also a major concern.
-
-
The most basic test SemiAnalysis applies to a GPU cluster health check is whether it could possibly work on paper. Several could not. "We've been through a few of these clusters where there's no way it could possibly work. These are health checks so bad they're worse than no health checks, because they're actively interfering with jobs." "On Amazon HyperPod Slurm, the first time we tested, there was a health check that needed the node to be healthy before it could run. It could only run once the node was back in the fleet, and it was needed to bring the node back into the fleet." "There are probably five or ten examples of this throughout testing where the health check never had a chance. It makes you wonder whether these people are really putting it through the paces before they hand it off to customers."
-
Singapore, you won't want to miss this: we're hosting the SemiAnalysis × Nanyang Capital Stock Pitch Challenge. Pitch an AI or semis stock for a shot at S$5,000 in prizes and a fast track into SemiAnalysis. Register by 11 Oct: https://lnkd.in/edjNeZTD
-
-
How GLM5.3 Sparse Attention Affects HBM Memory Usage GLM-5.3, KV Cache Offloading, HiSparse, AgentX TileRT, InferenceX DeepSeek Sparse Attention, IndexShare, Single-rollout Asynchronous Optimization, Cybersecurity https://lnkd.in/eBpQ3zXk