roomtsc

Model

roomtsc 0.1 takes a crystal structure and estimates its Hopfield sum, the quantity that sets the upper scale of a phonon-mediated transition temperature. For a hydride it also estimates the scattering strength per proton, h, which is the hydrogen part of that sum divided by the hydrogen density. It is trained on the Alexandria electron-phonon release at one atmosphere: 27,151 compounds with a spectral function, 4,619 of them hydrides with a separable hydrogen part. The first paper explains why these are the quantities to estimate, and the companion paper reports the model and its tests in full.

The estimate is the mean of two estimators, gradient-boosted trees on descriptors of the cell and an ensemble of graph networks. It is not a transition temperature. The labels are harmonic calculations at one smearing width, and nothing here says whether a structure can be made or kept.

Tests

The splits, the reference values and the baseline predictions were deposited before each run of the tests. Two of the tests the lab fixed in advance could be run on this data, with two cross-validations beside them. The others need calculations at megabar pressure or a reference set that does not exist yet.

The tests were run twice. The splits of the first deposit leaked: a structure type could carry two labels, so some compounds of a held-out type stayed in training, and five held-out compounds had a second record there. The second deposit labels structure types independently of the origin and keeps one record per compound. The table gives the second run. The first run's results are kept, with a measure of the leak.

Largest values held outCubic A₂MH₆ held outGrouped foldsRandom folds
Hydrides scored530724,5524,552
Electron-gas constant1.231.271.621.62
Constant × density of states1.171.171.061.06
Training mean0.740.830.630.62
Trees0.40 (0.53)0.33 (0.40)0.32 (0.51)0.22 (0.35)
Graph network0.32 (0.43)0.39 (0.47)0.34 (0.54)0.26 (0.41)
Mean absolute error in ln h on hydrides held out of training, on the splits of the second deposit. The first two tests hold out whole structure types: those with the largest values, and one family. Grouped folds assign whole structure types to a fold; random folds do not and carry no mark. For the two estimators the bracket is the ratio to the best of the three baselines, and the mark fixed before training is a ratio of at most 0.50. Full results with intervals: top, family, grouped, random.

The first test also asked whether the estimators predict above the largest value they were trained on. 23 of the held-out compounds exceed the training maximum of 34 eV Å. The trees place 0% of them above it and the graph network 9%. The requirement was at least half, so that test is failed. The estimators rank compounds and do not extrapolate upward.

The whole Hopfield sum, which the graph network also returns, was scored without a mark. Over the grouped folds its error in ln S is 0.22 against 1.27 for the training mean (26,058 compounds).

First use: the records without a label

84% of the release's hydride records have no spectral function, most of them because the structure is unstable in a harmonic calculation. The first paper names them as its largest gap, since strong coupling drives phonons toward instability. The released model was run on all of them: 30,822 compounds with an element besides hydrogen and a hydrogen density inside the range the model was trained on.

The estimates of h have a geometric mean of 10.7 eV Å, against 7.4 for out-of-sample estimates of the labelled hydrides. No estimate of the hydrogen Hopfield parameter reaches the 4.83 eV/Ų below which 300 K is out of reach; 3 exceed 3 and 112 exceed 2.

CompoundρH (Å⁻³)Lowest frequency (cm⁻¹)h (eV Å)ηH (eV/Ų)
H6KPt0.065−16734.8
H36Ag40.145−676263.8
H6NaPt0.067−50553.6
H6CuRb0.073−1366392.8
H6LiPt0.067−185422.8
H7HgN20.120−1075232.8
H7HgLi20.095−887292.8
H7AuBe20.115−730242.7
The eight largest estimates of the hydrogen Hopfield parameter, with the hydrogen density, the lowest harmonic frequency stored for the record, and the estimate of h. Formulas are printed as the release stores them. Full list: unlabelled_hydrides.csv, with a summary.

These are estimates and settle nothing about the gap. The model was trained on stable structures only, it failed the test of predicting above its training range, and for the platinum hexahydrides at the top it has labelled members of the same family in training. The list gives an order in which to calculate the Hopfield sum directly. The first candidates are the platinum hexahydrides whose instability is small.

Estimate a structure

Paste a structure as CIF or POSCAR, with every site fully occupied and at most 200 atoms.

Examples: KPtH₆, the largest hydrogen coupling in the release · Mg₂IrH₆, a published candidate · PdH, a known superconductor at 9 K

Weights and code

Weights: feature_order.json, graph_0.pt, graph_1.pt, graph_2.pt, meta.json, training_index.json, trees.joblib. The code that trains, tests and runs the estimators is with the other scripts on the data page (the files under model/). The weights load with the versions named in the deployment file there.

API

The form calls one endpoint, which takes the structure as text and returns JSON.

curl -s https://roomtsc.com/api/predict \
  -H 'Content-Type: application/json' \
  -d '{"structure": "…CIF or POSCAR text…"}'