LensVLM: Compressing long context as images, expanding only relevant pages
LensVLM is a model that compresses long context as images, expanding only relevant pages. It was released by Apple and has 9 billion parameters. The model is based on the VLM (Vision and Language Model) architecture, which is designed to process both visual and textual data. The model can be used for various applications, including chatting, question-answering, and text summarization. The article is available on the Hugging Face model hub.
Read the full article at huggingface.co →