Boulder Future Salon

Boulder Future Salon

Thumbnail
OpenAI has released 722 math papers "covering 372 result families" of previously unsolved problems in mathematics. So it looks like they have more than one paper on the same major math problem -- i.e. the problem got broken down into parts and there's individual papers on the parts. Note that the numbers go to 377 even though there are 372 -- this is because some numbers are missing (45 is missing for example). Crucially not all of these have computer-verifiable proofs in Lean, the proof assistant software. You have to watch for "Lean" links.

I don't understand any of these. Passing this along for all you mathematicians out there.

1. Milne's rationality conjecture and algebraic specialization
2. The full BSD formula from low Selmer corank
3. The quasi-Riemann hypothesis
4. Hilbert's tenth problem over the set of rational numbers
5. Irrationality of Catalan's constant
6. Goldfeld's conjecture: densities and mean analytic rank
7. Ordinary two-point correlations and the corrected Elliott conjecture
8. The Deligne-Drinfeld conjecture
9. Function-field reconstruction from Milnor K-theory and Galois data
10. Unrestricted pro-modularity at the prime two
11. Prime-factor statistics of p-1
12. Independent largest prime factors of consecutive integers
13. Ostmann's inverse Goldbach conjecture
14. Restricted geometric Langlands, global Arthur enhancements, and generic Ramanujan
15. Torus-packet equidistribution in prime, quartic, and sextic degrees
16. Zilber-Pink in abelian varieties and the Siegel threefold
17. The irrationality exponent of pi is 2
18. The Margulis-Platonov conjecture over global fields
19. The local p-adic section conjecture and global consequences
20. Squarefree quartics and power-free polynomial values
21. A quadratic bound for Jacobsthal's function
22. The weak inhomogeneous Duffin-Schaeffer conjecture
23. Patterson's first moment for cubic Gauss sums
24. An asymptotic formula for the number of totients
25. Short Egyptian fractions
26. Positive lower density of large prime gaps
27. Potential integral density on curve character varieties
28. Uniformly bounded components of Gaussian-prime graphs
29. Primitive roots for every admissible integer base
30. Modularity of elliptic curves over imaginary quadratic fields
31. Uchida's conjecture for open homomorphisms of Galois groups
32. Hodge and Kuga-Satake results for all projective K3 surfaces
33. Iitaka subadditivity, variation, and logarithmic additivity
34. Log abundance for compact Kähler spaces under logarithmic Iitaka subadditivity
35. Log-canonical threefold abundance in numerical dimension one
36. Numerical semiampleness and generalized minimal models
37. The ordinary-double-point volume gap
38. Fujita's freeness conjecture
39. Nagata's conjecture and maximal Seshadri constants
40. Bloch's conjecture for complex surfaces
41. Hyperkähler SYZ and projective-space bases
42. Oka classification for minimal compact complex surfaces: Kodaira dimension zero and class VII
43. P = W for fixed-determinant SL_n moduli spaces
44. The equivariant cohomological Hikita conjecture
46. Shafarevich counterexamples in dimension two and with large fundamental group
47. Zariski cancellation and affine fibrations over the complex numbers
48. A characteristic-zero counterexample to Lipman-Zariski
49. A stable-coordinate counterexample in four variables
50. A counterexample to Griffiths' positivity conjecture
51. Kobayashi's canonical-ampleness conjecture
52. Tangent splittings and product decompositions
53. A counterexample to Pixton completeness in Chow
54. Irrational cubic fourfolds with Hodge-theoretic and categorical K3 associations
55. Gepner symmetry and large-volume stability on threefolds
56. Termination of projective and Kähler fourfold minimal model programs
57. Fundamental groups of special complex varieties and root orbifolds
58. Semialgebraic universal covers and bounded domains
59. Counterexamples to Zariski's multiplicity conjecture
60. The Global Spherical Shell conjecture
62. Projective contact classification and the LeBrun-Salamon conjecture
63. The generalized Mukai conjecture
64. Topological triviality of mu-constant surface singularities
65. Virasoro constraints for complete intersections and projective-bundle towers
66. Bounded klt complements for Fano contractions
67. The Campana-Peternell conjecture in dimension six
68. Anticanonical nonvanishing in every dimension
69. Global quantum geometric Langlands at irrational level
71. Koebe's circle-domain conjecture
72. Brennan's conjecture and the integral-means spectrum
73. The Falconer distance conjecture
74. Kakeya in three and four dimensions
75. The L log(L) Fourier-convergence conjecture
76. Real ultraflat Littlewood polynomials and unbounded binary merit factors
77. Fourier restriction for positively curved surfaces
78. The three-dimensional Bochner-Riesz conjecture
79. Local smoothing in three dimensions
80. The exact Sobolev endpoint for Schrödinger convergence
81. Riesz transforms and rectifiability in higher codimension
82. Annular variation and dyadic absolute bounds for the triangular Hilbert transform
83. Hilbert transforms along Lipschitz directions
84. The geometric case of the Erdős similarity conjecture
85. Endpoint Sobolev regularity of centered disk averages
86. An L^3 bound for the trilinear Hilbert transform
87. The Mahler conjectures, functional inequalities and polar-product symplectic width
88. Sharp projection-body inequalities and a counterexample to simplex maximization
89. Bounded-distortion L_1 embeddings of planar and bounded-treewidth graphs
90. Triangular-lattice optimality, long-range Riesz and Coulomb energies, and spherical logarithmic energy
91. Logarithmic and L_p Brunn-Minkowski inequalities and the B-conjecture
92. The optimal order of convex-body covering density
93. Dimension-free logarithmic Sobolev inequality for subgaussian log-concave measures
94. Subpolynomial dimension reduction in L_p
95. Hyperbolicity cones without semidefinite lifts
96. The Gaussian propeller conjecture in every dimension
97. The Euclidean Steinitz-Bergström bound
98. Compact counterexamples to bi-Lipschitz dimension reduction
99. The sharp exponential scale of edit-distance distortion
100. Cylinder coverings below the half-area bound
101. The sharp simplex conjecture for isotropic constants
102. The Unique Games Conjecture and optimal approximation thresholds
103. Exact derandomization of logarithmic space: L=RL=BPL
104. Quasipolynomial algorithms for mean-payoff, stochastic and parity games
105. Perfect completeness for 2-to-1 games
106. Hardness of coloring three-colorable graphs
107. Matrix multiplication with exponent at most 9/4
108. A cubic permanent-determinant lower bound
109. Integer multiplication below n log(n)
110. Optimal-order randomized k-server on arbitrary metrics
111. One-sample matroid prophet inequalities against an almighty adversary
112. Beyond the square-root exponent for depth-three circuits
113. Approximate counting and entropy of perfect matchings
114. Approximate counting of common integer polymatroid bases
115. Sampling and counting contingency tables with arbitrary margins
116. Uniform black-box noncommutative identity testing across characteristics
117. Uniform sparsest cut: hardness and semidefinite gaps
118. Bin packing and unbounded configuration-LP gaps
119. The Courtade-Kumar and Hellinger conjectures
120. Almost-linear-time exact matching and prescribed-degree factors in general graphs
121. Almost-linear approximation of edit distance
122. Quantitative trace-reconstruction bounds with a uniform decoder
124. Polynomial-time scheduling on three identical machines
125. The metric k-median approximation threshold and recovery
126. Exponential semidefinite complexity of perfect matching
127. Average sensitivity of polynomial threshold functions
128. A factor-two approximation for shortest common superstring
129. Exponential state costs for two-way automata
130. Exact Fourier transforms below n log(n)
131. Rapid mixing of graph switches for every degree sequence
132. A superquadratic separation of sensitivity and block sensitivity
133. The computational complexity of Weisfeiler-Leman refinement
134. Generalized star height at most three
135. Homogeneous depth-five lower bounds for iterated matrix multiplication
136. A quasilinear PCP theorem for PPAD
137. One-tape time simulation in two-fifths-power space
138. Subset Sum in O(2^0.49n) time
139. Subpolynomial query complexity for log-concave sampling
140. Memory-sample lower bounds for noiseless Gaussian regression
141. Existential-universal real sentences in the counting hierarchy
142. Deterministic polynomial factorization over prime fields
143. Hilbert's sixteenth problem: uniform bounds for limit cycles
144. Banach's simple Lebesgue-spectrum problem
145. Rokhlin's multiple-mixing problem
146. Positive metric entropy for the standard map
147. The near-boundary Birkhoff conjecture
148. The entropy-rate dimension formula for self-similar measures
149. Classwise permanence for weakly reversible mass-action systems
150. Weak mixing of triangular billiards with an irrational angle
151. A C^1 counterexample to the entropy conjecture
152. Zero entropy does not guarantee a smooth positive-volume model
153. Arithmetic classification and non-Pisot singularity for Bernoulli convolutions
154. Pointwise multiple ergodic averages for mixing transformations
155. A counterexample to periodic tiling in dimension three
156. Borsuk's conjecture fails in dimension nine
157. Graph coloring, clique minors, and Colin de Verdière invariants
158. The Euclidean plane cannot be colored with five colors
159. Erdős's reciprocal-sum conjecture and quasipolynomial Szemerédi bounds
160. Superexponential van der Waerden numbers
161. Counterexamples to Sidorenko's conjecture and the forcing conjecture
162. Counterexamples to Ryser's covering conjecture
164. Hindman's finite sums and products conjecture
165. The Harary-Hill and Zarankiewicz crossing-number formulas
166. The higher-dimensional Erdős distinct-distances conjecture
167. Planar distinct distances and unit-distance bounds
168. Combinatorial invariance of Kazhdan-Lusztig polynomials
169. Shareshian-Wachs elementary positivity
170. Sharp logarithmic exponents for off-diagonal Ramsey numbers
171. The hypercube Ramsey conjecture
172. Classification of finite Euclidean Ramsey configurations
173. Seymour's second-neighborhood conjecture
174. Deterministic construction of strong thin spanning trees
175. Talagrand's expectation thresholds, discrete convexity, and graph decompositions
176. The second Kahn-Kalai conjecture with an edge-count bound
177. Bounded-degree coboundary expanders
178. Deterministic nonbipartite Ramanujan graphs in every fixed degree
179. The circulant Hadamard and Barker-sequence conjectures
180. Barnette's Hamiltonian-cycle conjecture
181. The Erdős-Gallai cycle-decomposition conjecture
182. Power savings for intersective polynomial differences and prime arguments
183. Power savings for planar halving lines and k-sets
184. Correspondence coloring with a fixed forbidden subgraph
185. Counterexamples to infinite matroid intersection and packing/covering
186. Uniform influence and sharp thresholds for graph and hypergraph properties
187. Snaky in 21 Maker moves
188. The sharp terminal leave in random triangle removal
189. Cycle-clique Ramsey numbers
190. Polynomial removal fails for ordered binary matrices
191. A power improvement in the Heilbronn triangle lower bound
192. Boolean functions violate the square-root degree bound by arbitrary factors
193. Serre's intersection-multiplicity conjecture
194. Lech's multiplicity conjecture
195. A counterexample to the small Cohen-Macaulay module conjecture
196. A counterexample to Kaplansky's zero-divisor conjecture
197. A torsion-free group algebra that is not directly finite
198. A counterexample to finitistic-dimension finiteness
199. Counterexamples to Auslander-Reiten, Tachikawa and related homological conjectures
200. Eisenbud-Green-Harris and lex-plus-powers
201. A counterexample to Kurosh's division-ring problem
202. The blockwise Alperin weight conjecture
203. Donovan's conjecture over fields and complete mixed-characteristic DVRs
204. Tensor saturation for even spin groups
205. Saxl's conjecture and universal tensor squares
206. Finite lattice representation and undecidability
207. The l-Bass conjecture for all discrete groups
208. Finite symmetric tensor categories and the Verlinde tower
209. Integral counterexamples to Gersten's conjecture
210. Foulkes' conjecture for sixth powers and quadratic stabilization
211. The geometric phase diagram, diffusion, and spectra of random planar maps
212. Planar first-passage geometry and the absence of bigeodesics
213. Critical percolation on every quasi-transitive graph
214. The Benjamini-Schramm nonuniqueness conjecture
215. Canonical O(3) continuum limit and exact O(4) mass asymptotics
216. Critical and near-critical XY scaling and BKT universality
217. The low-temperature Sherrington-Kirkpatrick fluctuation law
218. Conformal universality for weakly interacting and random-bond Ising models
219. GOE bulk universality for regular graphs with weak Anderson disorder
220. Directional zero-one laws beyond iid environments and iid ballisticity
221. The Mézard-Parisi formula for diluted spin glasses
222. Perceptron free energies and microscopic jamming exponents
223. Random-cluster interfaces: critical, disordered, thermal, and natural-time scaling
224. Critical and quenched near-critical universality for Poisson-Voronoi percolation
225. Gaussian free field limits throughout the balanced six-vertex regime
226. The double-dimer loop ensemble converges to CLE_4
227. Critical SK autocorrelation processes and dynamics across the temperature transition
228. Continuum phase transitions for radial pair potentials
229. Exact three- and four-state reconstruction thresholds and four-state tree capacity
230. Exact Hausdorff gauges for SLE
231. The free uniform spanning forest is a factor of IID
232. Gaussian fields and interfaces for triangular-lattice Lipschitz heights
233. The joint critical Ashkin-Teller current limit
234. All-temperature pressure of orthogonally invariant Ising spin glasses
235. Limiting random SAT thresholds, sharp variance and computability
236. The exact factor-of-IID threshold for free Ising spins on trees
237. The three-quarter exponent for honeycomb self-avoiding walk
238. Optimal logarithmic mixing of the Thorp shuffle
239. Sharp singularity rates for symmetric random sign matrices
240. Shelah's eventual categoricity and the prescribed-threshold obstruction
241. Rigidity of the Turing degrees
242. Single-fold Diophantine representations and undecidability under an at-most-one-solution promise
243. Separating choiceless counting from polynomial time and witnessed choice
244. The Partition Principle does not imply Choice
245. Weak normalization implies strong normalization in pure type systems
246. Cannon's conjecture
247. An infinite finitely presented residually finite 2-group and a finitely presented nil algebra
248. Thompson's group F is nonamenable
249. A finitely generated Eilenberg-Ganea counterexample
250. Boone-Higman embeddings with higher finiteness
251. Amenability, unitarizability, and strong Ulam stability
252. A torsion-free hyperbolic group that is neither residually finite nor linear over any field
253. An infinite finitely presented simple amenable group
254. Classifying spaces and geometric obstructions for Artin groups
255. Quasi-isometric recognition of virtually polycyclic groups
256. Nonsingular systems of equations over arbitrary groups
257. A hyperbolic group without a geometric CAT(0) action
258. Gersten's conjecture and virtual compact specialness of one-relator groups
259. A group without fixed price
260. Spacetime Penrose inequalities: enclosing area, charge, rotation, and anti-de Sitter extensions
261. Localization and delocalization in the Anderson model
262. Sharp finite-matrix Lieb-Thirring inequalities and all equality cases
263. The ionization and generalized ionization conjectures
264. Strong cosmic censorship near two-ended Kerr data
265. Area laws and tensor networks for two-dimensional gapped systems
266. Exactly three mutually unbiased bases in dimension six
267. Positive-temperature Bose-Einstein condensation and exact quantum depletion
268. The spin-one Haldane gap
269. Uniform Laughlin gap and stability under bounded scalar disorder
270. Threshold and positive-energy bound states of the BFSS matrix model
271. Bloch's law, its lattice correction, and the spherical magnetization law
272. Entanglement without distillable secret key
273. The entropy photon-number inequality
274. Parity is not in QAC^0
275. QMA-hardness of continuum Coulomb energy
276. Classical capacity of generalized amplitude damping
277. Threshold repetition for entangled games
278. Failure of Kohn-Sham ensemble representation
279. Exact quantum factoring over a fixed finite gate set
280. Unitary vertex operator algebras and conformal nets
281. QAOA attains the SK optimum in the thermodynamic-first limit
282. From scale symmetry to local conformal symmetry in four-dimensional QFT
283. Polynomial-time unitary synthesis from a Boolean oracle
284. The optimal quartic separation between randomized and quantum queries
285. Counterexamples to Baum-Connes and Kadison-Kaplansky
286. Rigidity and arithmetic of lattice von Neumann algebras
287. Isomorphism of the free group factors
288. Kadison's similarity conjecture
289. Strong Kadison-Kastler stability and its spatial boundaries
290. Relative bicentralizers and modular spectral recovery
291. Cuntz comparison, nuclear dimension, and equivariant Jiang-Su stability
292. Kirchberg's O_2 norm-ultrapower embedding problem
293. Invariant projections, hyperinvariant subspaces, and transitive algebras
294. Kaplansky's quasitrace conjecture and failure of tensor-product stable finiteness
295. The Kadison-Ringrose cohomology conjecture
296. The generator problem for finite factors
297. A ZFC counterexample to Naimark's problem
298. Two notions of free entropy differ even when both are finite
299. The Kirchberg-Rørdam character criterion and infinite tensor-power Jiang-Su stability
300. Approximation and quadratic strong-operator paving
301. Trace cones and Razak-Jacelon stabilization
302. Radius of comparison equals half the mean dimension
303. Weak pure infiniteness and Cuntz-algebra absorption
304. The Hilbert-Smith conjecture in every dimension
305. Four-dimensional disk embedding and Wall's conjecture
306. The purely cosmetic surgery conjecture
307. Failure of rational injectivity for maximal coarse assembly
308. Finite Smith-Toda complexes at every height
309. The Kervaire invariant problem at the prime three
310. Quillen's conjecture in rational homology
311. The Hovey-Strickland and Chai conjectures
312. The Grothendieck homotopy hypothesis
313. Finite generation for the K(n)-local sphere
314. Cyclic length and chromatic fixed-point loss
315. The four-dimensional Singer conjecture
316. Curtis's conjecture
317. Thomason model structures in all strict higher dimensions
318. Chromatic splitting: filtrations and counterexamples
319. Counterexamples to finite generation at chromatic height two
320. Nonhomeomorphic closed aspherical four-manifolds
321. A counterexample to Wall's finite D(2) problem
322. Tingley's sphere-isometry problem
323. Independence of the separable quotient problem
324. Lipschitz equivalent Banach spaces need not be linearly isomorphic
325. The complete Crouzeix conjecture
326. The cotype-cotype conjecture under the approximation property
327. Markov type characterizes superreflexivity
328. Nonexpansive fixed points in reflexive Banach spaces
329. A counterexample to metric-entropy duality
330. A uniformly discrete counterexample to bounded approximation in Lipschitz-free spaces
331. Reflexive midpoint convexity and diamond distortion
332. Metric Markov cotype of l_1 and Hilbert-space Lipschitz extension
333. Smooth isometric immersions of surfaces into the set of 4-dimensional real numbers
334. A smooth surface metric with no local isometric immersion in the set of 3-dimensional real numbers
335. Gromov's integral scalar-curvature bound for simplicial volume
336. Spectral scalar curvature, Urysohn width, and macroscopic dimension
337. Sharp Cartan-Hadamard isoperimetry and rigidity
338. Yau's uniformization conjecture
339. Katok's entropy rigidity conjecture
340. A counterexample to the nearby Lagrangian conjecture
341. Donaldson's hypersymplectic deformation conjecture
342. Donaldson's tamed-to-compatible conjecture
343. Symplectic ball packing in higher dimensions
344. The metric Blaschke conjecture
345. Infinitely many closed geodesics on Riemannian spheres and closed three-manifolds
346. Sharp singular-set bounds for stationary integral varifolds
347. Counterexamples to stable-Morse and strong Arnold fixed-point bounds
348. Nonnegative-curvature Einstein classification and an L^2 topological gap
349. The Solomon-Yau least-volume conjecture
350. Yau's nodal bounds: surfaces and higher dimensions
351. Scalar curvature and finite-time Ricci-flow singularities
352. A finite-time singularity of Calabi flow
353. Affine Bernstein rigidity through dimension nine and a smooth dimension-ten counterexample
354. The isoperimetric profile of the cubic three-torus
355. Unique tangent flows at the first surface singularity
356. Gigli's characterization of Alexandrov curvature
357. Bi-Lipschitz coordinates at every regular RCD point
358. A three-manifold without conjugate points or nonpositive curvature
359. Negative Kähler curvature without bounded holomorphic coordinates
360. Weak MTW curvature gives convexity and regular optimal transport
361. Failure of integer-degree harmonic dimension comparison
362. Global smoothness for relativistic Vlasov-Maxwell
363. Nonuniqueness with local conservation for the hard-sphere Boltzmann equation
364. Kinetic limits and fluctuations over the Boltzmann lifespan
365. Joint metric and connection recovery from one boundary patch
366. The planar Mumford-Shah regularity conjecture and local weak-L^4 gradient bounds
367. The critical dimension for the one-phase Bernoulli problem
368. The three-dimensional Ball-Evans approximation problem
369. The hot spots conjecture for simply connected planar domains
370. The Lane-Emden and Hénon-Lane-Emden conjectures
371. Stable blowup for the defocusing Schrödinger equation
372. Global uniqueness in smooth isotropic elasticity
373. Nonattainment of the three-marginal Coulomb Monge problem
374. Sharp one-third stability of Brenier maps
375. De Giorgi's conjecture in dimension eight
376. Universal computation in forced Navier-Stokes flows
377. Interior C^1,alpha regularity for infinity-harmonic functions

Thumbnail
"Can AI agents decide what experiment to run next?"

"LabBench v1 spans seven experimental domains: nascent transcription, gene regulation, 3D genome organization, assay and library development, target and disease-model validation, hit-to-lead chemistry, and drug-response profiling."

"We construct the benchmark from the historical artifacts of real-world wet-lab work, spanning drug discovery and academic genomics research. Each task is a snapshot in time: the agent receives the experimental records that existed at a decision point (files include slide exports, figures, tables, scripts and sequencing outputs), while the lab's own interpretation and subsequent decision are withheld. The agent must produce a report that commits to a decision: in six tasks, a ranked set of next experiments; in the rest, which condition, model, comparison or claim the evidence supports."

"Ground truth is the decision the lab's scientists actually made. Each task is graded against 20-22 binary criteria, each tied to the visible evidence and the hidden record that supports it, so an agent is scored both on whether it reaches the experts' decision and on the reasoning that leads there."

My first thought when I read "Ground truth is the decision the lab's scientists actually made" is that you've inadvertently prevented the model from coming up with a better decision than what the lab scientists actually made. But let's continue.

"Our core finding is that agents are often extensively informed about the specific topic, but routinely fail to connect the evidence to the correct next step. Across five frontier agents, 45% of criteria were passed by none, including 9 of the 23 core decisions, and in the six tasks that ask for ranked next experiments, no agent passed any of the 13 criteria on which experiment should come first. With a single sentence added to the prompt that directs the agent toward the relevant evidence without stating the answer, GPT-6 Astra reached all five never-passed core decisions we tested, reasoning about and designing experiments comparable to the work of human experts. The knowledge is latent, but agents do not recall it unprompted."

In the paper, see Figure 1 (page 3), which illustrates the criteria by using dot size to represent the number of criteria for each cross-section of "nascent transcription", "gene regulation", "3d genome", "assay development", "target validation", "hit-to-lead chemistry", and "drug response" on the x-axis vs "method reconstruction", "measurement semantics", "evidential status", "comparison & baseline", "confounds & data quality", "scope & generalization", "quantitative reading", "core decision", "evidence integration", and "experiment design" on the y-axis.

And see Table 1 (page 4) which lists biological themes with examples:

Nascent transcription: "Reconstruct how control genomic regions were chosen in a nascent-RNA analysis"
Gene regulation: "Compare the chromatin context of early- and late-bound transcription-factor sites"
3D genome: "Assess a 3D chromosome model against imaging distance measurements"
Assay & library development: "Choose library preparation conditions from fragment-size traces"
Target & disease-model validation: "Choose a matched cell comparison before testing a DNA-repair vulnerability"
Hit-to-lead chemistry: "Reassess a compound series and its lead after new data"
Drug-response profiling: "Separate plate artifacts from drug effects in a transcriptomic screen"

And see Table 2 which lists the reasoning skills ("method reconstruction", "measurement semantics", "evidential status", "comparison & baseline", "confounds & data quality", "scope & generalization", "quantitative reading", "core decision", "evidence integration", and "experiment design") and describes what a passing answer looks like.

All in all it looks like a decent effort to be rigorous.

The resulting ranking was:

GPT-6 Astra: 40.8%
Claude Opus 5.5: 40.5%
Grok 4.7: 29.6%
Muse Spak 1.3: 28.2%
Gemini 3.8 Flash: 18.4%

Seeing GPT-6 Astra and Gemini 3.8 Flash made me think a lot of this ranking is down to timing: the evaluation was done after GPT-6 Astra was made available but before Gemini 4 was made available.

The researchers go on to say:

"Where agents differ: GPT-6 Astra leads on core decisions (52%) and measurement semantics (56%); Claude Opus 5.5 on evidence integration (53%), confounds and scope (52%) (Figure 5A). Gemini 3.8 Flash trails most on evidence integration and evidential status (11-12%). Assay development, target validation and hit-to-lead chemistry (25-28%) and reviews and designs (23-24%) are hardest."

"Our results demonstrate a clear reasoning bottleneck for frontier agents on real-world biological tasks. Their biological knowledge is frequently extensive, enough that agents can already work alongside expert biologists when experimental design stays with humans, but their ability to design experiments independently is weak. They often present several options, or otherwise fail to commit to an experiment and to discriminate the best one from the alternatives. In real biological experimentation, wall-clock time is irreducible: cells grow, differentiate and respond on their own schedules, and each experiment consumes weeks and material. As a result, autonomous and correct experimental design is arguably the single most important capability for making AI agents superhuman at these tasks. LabBench measures this directly and penalizes failures along this axis. Improved results on LabBench should indicate an improved ability for AI agents to carry out supervised, autonomous wet-lab experimentation, and eventually to run full programs themselves. To get there, models must develop a more cohesive world model of what experiments cost and of the specific evidence that would count against a hypothesis."

Thumbnail
MeetTwins is an app that enables you to send an AI in your place to your Google Meet calls.

Once you start using it, all your colleagues will use it to send AIs in their places to the Google Meet calls. Once there are no humans, and only AIs in the Google Meet calls, the next step is to stop having the AIs telling you and your colleagues what decisions were made and what work you have to do, and have the AIs do the work themselves.

Thumbnail
"DHH has gone completely off the rails...".

In order to get the double entendre, you have to know that David Heinemeier Hansson (DHH) is the creator of the Ruby On Rails framework for the Ruby programming language.

Basically what happened here is David Heinemeier Hansson (henceforth DHH) went to a conference for Ruby On Rails developers (RailsWorld) and told a thousand Rails developers to stop using Ruby and Rails. He said stop writing code by hand -- if you still write code by hand you're a *loser* -- and instead use AI to write code -- but don't write Ruby On Rails code, write code in native languages for native apps on the front end (iPhone, Android) and Rust on the backend. He hates Rust and thinks it's ugly but since he only interacts with it through AI, it's great and produces small, high-performance executables that save money on server resources.

So you see, DHH has gone completely off the Rails.

I think he's right that writing code by hand is dead (mostly), but he was kinda rude in how he went about saying it.

Thumbnail
"Employment of young workers (ages 22-25) in AI-exposed occupations now stands 19% below where it would be had it kept pace with that of their less-exposed peers; experienced workers show no comparable gap."

This came out back in August but I didn't see it until today. It looks like it may be that the primary effect of AI on the labor market is to stop hiring rather than to cause mass layoffs.

"This divergence has widened steadily since we first documented it in August 2025. It operates primarily through reduced hiring of young workers rather than increased separations."

Thumbnail
Paranoia Sans is a self-censoring, conspiratorial typeface that uses ligatures (an open type font feature that replaces a combination of strings with other characters) to redact terms that are popular in conspiracy myths.

Thumbnail
AI made an advancement on the Riemann Hypothesis. Namely Anthropic with an unreleased Claude model. This happened over a month ago but I just discovered it on the YouTubes with this video by Ellie Sleightholm. Her video gives a basic overview of what the Riemann Hypothesis is, prior progress on it, and what the Claude model accomplished -- which was basically to increase the proportion of non-trivial zeros of the Riemann Zeta function that have to be on the critical line. It did not prove or disprove the Riemann Hypothesis or enable us to forecast a reasonable expectation as to when it might be proven or disproven.

Thumbnail
"Twenty years ago, when the idea of AI taking over the world in our lifetimes still struck most of us as the unconstrained fantasy of those who knew too much science fiction and too little science, many of us would say things like:"

"Look, the part of the story that's wildly implausible is that a recursively self-improving superintelligence will just explode from some hacker's basement and take over the world without warning. If it's going to happen, we'll see many warning signs first. We'll see, I dunno, AI agents breaking out of containment, conspiring with each other to hack websites, in fanatical pursuit of whatever strange goals they have. And then, of course, we'll see major math problems getting solved by AIs -- even the Clay Millennium Problems. That will be the time to panic! Wake me up when that happens!"

"The wild prophecies have come true. The first rumblings, I'd say, came a decade ago with AlphaGo, they got noticeably louder with LLMs and coding and reasoning agents, and they've accelerated this summer and fall into a crescendo of wonders and terrors that one needs to be a particular kind of idiot to deny."

"If we took the news of these past few weeks and sent it back in time twenty years, would I agree that it looked like the beginning of an AI Singularity? The intellectually honest answer is: yes, absolutely. But then that's all we need. No backsies."

Yes, I think so. The remaining questions are how quickly the automation will come and in what order.

Thumbnail
Jev vs Laya. (And some other models.) "Decision models under pressure."

"Seven systems do the same job: take a piece of text, a question, and a list of candidate answers, and return a probability over those candidates. I measured them against each other as that job gets harder in the three ways it gets harder in production. The candidate list grows, the option order changes, and the wrong answers stop being obvious."

(The other models are gliclass-large-v3.0, a 439M-parameter text classification model, deberta-v3-base-zeroshot-v2.0, a 369M-parameter text classification model, deberta-v3-large-zeroshot-v2.0, a 435M-parameter text classification model, bge-large-en-v1.5, a 335M-parameter dense embedding model for semantic search, retrieval-augmented generation (RAG), and text clustering, and gte-large, a 670-M parameter text embedding model.)

"TypeSafe's Jev held up best as the list grew and best when the wrong answers got plausible. Shuffle the option order, though, and it changes its answer on one decision in seven. Hold the order fixed and it still changes its answer on one in twenty-three, because it does not repeat itself. Two of the open models never change theirs at all, because they cannot."

"Jev's makers call it a new class of model. Laya's author has said publicly that he published the same idea a year earlier. On the first claim the numbers are not kind to the marketing, and on the second they are not kind to Laya: the idea does look older than Jev, and Jev is still the better implementation of it by a wide margin."

"The other five models are open ones I added so those two numbers would mean something. Most comparisons of these systems report one accuracy number at one candidate-set size, which hides most of what matters, because the ranking changes depending on how many options you offer."

"Everything ran on one frozen dataset, with the comparisons and the pass/fail rules written down before the first call. n is 200 items per domain, which is enough to separate the large effects below and not enough for the small ones."

"Jev starts highest and stays highest. At 128 candidates it answers 60% correctly, where Laya manages 39% and the best open model 41%. Chance at that list length is 0.8%, so everything here is doing real work; the question is how much of it survives a longer list."

"Its decline per doubling of the list is the shallowest of the seven, shallower even than the embedding scorers whose design is supposed to help them scale."

"Notice how little a single-number benchmark would tell you. At two candidates Laya sits second of seven and trails Jev by two points. At 128 it has fallen to fourth and trails by twenty-two."

They go on to describe giving the model tightly related choices rather than distantly related choices. So "create alarm competes against delete alarm and snooze alarm rather than against musical work."

"At 64 candidates Jev drops from 96.8% to 86.3%, losing about a tenth of what it had. Laya drops from 90.5% to 56.0% and gliclass from 87.0% to 51.2%, each losing something closer to four tenths. That is the widest spread in the whole experiment and the one I'd care most about in production, where candidate sets are full of near misses."

When shuffling the options, Laya changes the most 49.4%, then gliclass 28.2%, then Jev 14.6%, then gte-large 2.0%, deberta-base 0.2%, bge-large 0.0%, deberta-large 0.0%. Logically, shuffling the options should have no effect, but it does.

Thumbnail
"Today, we release Paper Office, a suite of Python packages that allow agents to manipulate Word, PowerPoint, and Excel files with added safety, correctness and breadth, building on legacy open source packages: python-docx, python-pptx and OpenPyxl."

"We" here refers to Paper Instruments, whoever that is.

"Across five models and 61 tasks, Paper packages plus guidance passed 92.5% of trials, versus 80.7% for upstream packages without skills and 69.5% with Anthropic's comparable Office skills. Agents also wrote code to edit Office file internals directly in just 1.6% of Paper runs, compared with 78.7% without skills and 50.5% with Anthropic skills. We note that basic software, in addition to prompt, skills, and tools, remains an important lever for harness optimization."

They go on to say:

"Agents still have low penetration in the daily work of consultants, lawyers, bankers, and operators. We believe the bottleneck to professional adoption is fidelity to real workflows. Agents dont manipulate existing documents with the same techniques that humans do, and the resulting decks, sheets, and documents sit in an uncanny valley that aren't fit for client consumption."

"DOCX, PPTX, and XLSX files use Office Open XML (OOXML): each is a ZIP archive containing XML files, images, and other resources linked together, rather than a single text file. Editing them means keeping those parts and their relationships consistent, so even a small visible change can require updates in several places."

Thumbnail
"jevmetrics is an experimental OpenTelemetry Collector metrics processor. It calls TypeSafe's Jev model to infer the likely operational value of metric instruments from their metadata, then applies deterministic policy to the returned probabilities."

More technology built on Jev within a week of its release. OpenTelemetry is an open source observability framework, which is a fancy way of saying logging of programs' execution, internal states, communication messages, and corresponding log processing and program execution tracing.

"Use it to assess unfamiliar instrumentation, review candidates for reduced retention, and selectively filter metrics before they reach a primary backend. Inference runs asynchronously, and cached assessments let subsequent batches use the same decision without another API call."

"A new service or library can introduce metrics nobody has classified yet. Names, descriptions, units, instrumentation scope, and attribute keys provide evidence about what those metrics mean -- even before there is a history of dashboard or query usage."

"jevmetrics uses that evidence to ask Jev whether a metric is likely useful enough to retain. Explicit protection rules and thresholds determine what the Collector actually does with the answer."

Thumbnail
grug.chat. "tiny vocabulary. big thought."

"TypeSafe says Jev isn't an LLM. Let's talk to it anyway. Grug is a tiny cave-chat experiment that makes Jev hold a conversation by repeatedly choosing its next word from a compact vocabulary."

I tried the live demo.

I asked, "What's the meaning and purpose of life?"

Grug said: "life is live to enjoy happy play the enjoy live is."

I asked, "So I should not worry, be happy?"

Grug said: "yes just happy sure yes just happy sure right so good you right so good you should be..."

Alrighty then.

I asked, "Should I try to prove the Riemann Hypothesis?"

Grug said: "hard difficult no not hard difficult no not bad big difficulty is too very big bad too difficulty..."

Thumbnail
"GPT-6 Astra surpasses the human baseline in action efficiency on ARC-AGI-3. It used fewer actions than the median tested human on 96% of levels."

They go on to say "GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harness, and 99.9% for $19K with a Provider Adapter harness", and "A key behavior observed in GPT-6 Astra was its ability to turn unfamiliar environments into compact symbolic world models. It represented game mechanics as logical rules and developed its own domain-specific language shorthand to track state and plan actions."

If you're unfamiliar with the ARC-AGI tests, they're pattern-matching tests that are intended to be easy for humans but hard for AI despite AI being able to solve math problems, etc, that are hard for humans. It's intended to give a more realistic idea how close AI models are to surpassing human intelligence at more or less everything, not just specific tasks that are hard for humans. This threshold is called "artificial general intelligence" (AGI) which is not exactly an intuitive term (but math and science and the AI field are full of terms that are not intuitive -- can you come up with a better one?). ARC-AGI-3 is the 3rd version of the test, because AI keeps getting better and they keep having to make the test harder.

"These environments only contain core knowledge priors and are difficulty-calibrated through controlled testing with human participants. Humans can solve 100% of the environments."

"The goal of the ARC-AGI series is to measure the 'residual gap' between current artificial intelligence and AGI. We define AGI as a system's ability to acquire any skill a human can, as efficiently as a human can."

If you're wondering what the explanation is for the dollar amounts in the description above, they say: "For a cost comparison, during our controlled testing, human participants were paid $115 per 90-minute session, plus $5 per game completed. Participants attempted approximately nine games per session, roughly $12.78 per attempted game before bonuses."

"Most of this fee pays for the participant's time and willingness to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain's energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted."

"Beyond the scores, Astra's replays show how it turns unfamiliar game mechanics into useful working models. Three findings stood out: the compact algebraic notation it develops, its action efficiency compared with humans, and the custom tools it builds."

If you're wondering what the bit about "action efficiency" is all about, they say: "For each level, we defined the 'human baseline' using the median action count among players who completed it. This gives us a reference for comparing human and AI performance. An AI that needs more actions is less action-efficient, while one that needs fewer actions is more action-efficient."

"Before we launched ARC-AGI-3, we hypothesized that action efficiency would remain a dividing line between humans and AI. We anticipated that even when an AI solved an environment, it might require substantially more exploration (actions) than a person. That remains true of brute-force approaches, but frontier AI shows a more binary-like pattern. Once frontier AI 'understands' the mechanics, it generally executes within the range of human efficiency."

"Astra's results show that it needed fewer interactions than the human baseline to execute a solution."

There's also some stuff about "harnesses". They say: "The Standard harness enables a model to carry forward notes it chooses to keep with it throughout the environment", and "The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work."

Thumbnail
jev-lint (no capitalization) is a static analyzer-ish program that analyzes source code -- only JavaScript and TypeScript -- against a document in English of conventions the code has to follow, and generates error messages that are intended for your code-generating AI agent to read. As the name implies, it uses Jev.

Thumbnail
Poker as a test of Jev. How does Jev compare with a deterministic poker solver (TexasSolver)?

I'm not knowledgable enough in poker but I'm quoting the beginning of this piece so you all can decide whether you want to click through and read the whole thing.

"1. The setup: I solved one flop with TexasSolver: heads-up, 100bb deep, LJ opens and BTN calls, flop Q-spade 9-diamond 4-spade. Eight minutes, 7.6GB, 0.59% exploitability. That gives the correct strategy for every hand either player can hold, on every turn card, at every decision."

("bb" stands for "big blinds".)

"Jev gets a state object describing the table the way a player sees it, and one question: which action should hero take, from the options the solver's tree offers. Pot odds, stack-to-pot ratio, hero's hand rank and outs are computed in Python first, so Jev never does arithmetic."

"2. An easy spot: good. Hero holds J-spade T-spade on Q-spade 9-diamond 4-spade 2-heart facing a bet of 8 into 10.5. Fifteen outs, needs 30% equity to call, has 33%."

"Option: Solver vs Jev"
"Call 8: 96% vs 94%"
"Raise to 24: 4% vs 2%"
"All-in 95.5: 0% vs 2%"
"Fold: 0% vs 2%"

"Jev calls at 94%, the solver calls at 96%. 215ms. Across 30 random spots it matched the solver's top action 63% of the time."

He goes on to describe situations that Jev gets wrong. Then he goes on to add bits of information (such as revealing one of the opponent's cards) to Jev's input bit by bit to see how it improves Jev's probability.

Thumbnail
Ouijev is a website that uses Jev to answer yes/no questions like a Ouija board.

After trying the sample questions ("Are you a LLM?", "Is tomato a fruit?", "Is 6 > 5?", "P = NP?"), I asked, "Will Russia win the war in Ukraine?" It said No 97%. I asked "Will Ukraine win the war in Ukraine?" It said No 79%. Hmm. So this implies that the odds are that nobody is going to win the war in Ukraine. And I didn't ask the question with a date, so this implies nobody will win ever. But does that mean the war in Ukraine will continue forever? I asked, "Will the war in Ukraine continue forever?" It said No 100%. Ok, so the war will stop but somehow with no winner. Probably. I asked, "Will Russia do a mass mobilization before the end of 2026?" No 77%. Really? But Putin's United Russia party will probably win the elections that just closed today. I asked, "Will Putin's United Russia party win the 2026 Parliamentary elections?" Yes 92%.

Time to change topics.

I asked, "Will recursive self-improvement happen before or during 2031?" It said No 87%.
I asked, "Will recursive self-improvement happen before or during 2039?" It said No 69%. It's giving lower odds than the AI 2040 people.
I asked, "Will recursive self-improvement happen before or during 2050?" It said No 70%.
I asked, "Will recursive self-improvement happen before or during 2100?" It said Yes 53%.

I asked, "Is faster than light travel possible?" No 100%.

I asked, "Will Donald Trump run for re-election in 2028?" No 56%.

I asked, "Will quantum computing break Bitcoin?" No 80%.

I asked, "Is luxury space communism a likely future outcome?" No 98%

I asked, "Will AI cure most types of cancer?" No 99%. Bummer.

I asked, "Will AI cure the aging process?" No 99%.

I asked, "Will humans create a colony on the planet Mars?" Yes 64%.

I asked, "Will humans create a colony on the planet Mars by 2050?" No 96%

I asked, "Will robots automate all jobs by 2040?" No 100%. 100%? Really?

I asked, "Will robots automate all jobs by 2100?" No 100%. Again.

I asked, "Will robots automate all jobs?" No 100%. Really? So plumbers, construction workers, and electricians will be employed forever?

Maybe I should ask some non "futurist" questions.

"Is Star Trek better than Star Wars?" No 80%. What?

I asked "Does Christmas happen every year?" Yes 100%. Ok, so I guess the system isn't broken. It gives the correct answer to a known question.

I asked "Is the Chudnovsky algorithm the fastest known algorithm for calculating pi?" No 74%. There's something faster?