The Roman Galaxy Redshift Survey will measure redshifts from slitless grism spectroscopy, and testing how reliably those redshifts can be extracted requires realistic simulated surveys - mock skies whose galaxy number densities, sizes, and emission-line strengths are right, not merely plausible. We build these by applying the Galacticus semi-analytic galaxy formation model to dark matter halo merger trees from the UNIT N-body simulations, and placing the resulting galaxies on a lightcone. This only works if Galacticus is calibrated against the observables that matter, and calibration is expensive: a single likelihood evaluation involving a stellar mass function requires evolving thousands of halos.
I will describe our approach to this problem. Building on an earlier fast- likelihood method that calibrates against the stellar-to-halo mass relation rather than the mass function directly, we now emulate Galacticus with Gaussian processes. From a 1024-point Latin hypercube spanning 20 model parameters, we train emulators — PCA-compressed, one Gaussian process per component — that predict stellar mass functions, star formation rate functions, sizes, and Hα luminosity functions in milliseconds, with calibrated uncertainties. Held-out tests show accuracies of 0.05–0.13 dex where the observational data lie. This makes simultaneous MCMC calibration against many datasets tractable, and reveals which observables constrain which parameters — including tensions between mass functions and galaxy sizes that point to genuine missing flexibility in the model physics rather than shortcomings of the fit. I will close with the status of the calibrated 16 deg² lightcone being delivered for the team's grism image simulations.
