I agree with your general statement, but if one thing the recent DeepMind Chinchilla research paper showed us is that the size of the model (number of parameters) is much less of the determinant of model performance & quality than the amount of high-quality data (number of high-quality tokens).
Their 70B chinchilla model significantly outperforms the 175B GPT3 model.
Possibly where OpenAI has a leg up is their high-quality data sourcing & curating infrastructure and their RLHF mechanisms.
Their 70B chinchilla model significantly outperforms the 175B GPT3 model.
Possibly where OpenAI has a leg up is their high-quality data sourcing & curating infrastructure and their RLHF mechanisms.
Paper: https://arxiv.org/abs/2203.15556