The ML part of CS is in a sense funding replication: there are a decent number of papers whose only premise is "we compare 5 recent papers on this benchmark (and fill a couple pages with discussion)" or "we made a new benchmark, here's how commonly cited papers compare (plus a couple pages of discussion)"
Outside the high-profile cases it seems accepted norm that papers perform far worse when scored against somebody else's benchmark. The real measure of quality is how big the gap is.
Outside the high-profile cases it seems accepted norm that papers perform far worse when scored against somebody else's benchmark. The real measure of quality is how big the gap is.