Mercury 2.5 LLM hits 770 tokens per second
The Mercury 2.5 large language model (LLM) has achieved a processing speed of 770 tokens per second, making it a significant improvement over its predecessors. The model is based on a transformer architecture and has 340 million parameters. Its performance on various benchmarks has been measured, with notable improvements in speed and accuracy. The Mercury 2.5 model is a development of the Mercury 2.0 model, which had a processing speed of 380 tokens per second. The new model has been trained on a larger dataset and has shown improved performance in tasks such as language translation and text summarization. The authors of the model suggest that it could be used for applications such as chatbots and virtual assistants.
Read the full article at artificialanalysis.ai →