Machine-learning run comparison

Compares experiment runs from the metrics you paste: validates the JSON, matches metric sets (a missing metric is a warning, not an error), computes minimum, maximum, mean and spread per metric, ranks runs by the primary metric in the max or min direction, flags dominated runs (no metric better than another run and at least one worse) and lists which parameters differ between runs. It works locally, without an MLflow server.

Fill in the fields, run the tool and review the result. Your input is not added to a public page. Use the learning mode for calculation details.