rw-book-cover

Metadata

Highlights

  • . Qwen is Alibaba Cloud’s organization training LLMs. Their latest model is Qwen2-VL - a vision LLM - and it’s getting some really positive buzz. Here’s a r/LocalLLaMA thread about the model. (View Highlight)
  • The original Qwen models were licensed under their custom Tongyi Qianwen license, but starting with Qwen2 on June 7th 2024 they switched to Apache 2.0, at least for their smaller models:

    While Qwen2-72B as well as its instruction-tuned models still uses the original Qianwen License, all other models, including Qwen2-0.5B, Qwen2-1.5B, Qwen2-7B, and Qwen2-57B-A14B, turn to adopt Apache 2.0 (View Highlight)

  • Here’s where things get odd: both of the above links are to the Internet Archive, because at some point in the last 24 hours the Qwen GitHub organization, and their GitHub pages hosted blog, both disappeared and are now 404s pages. I asked on Twitter but nobody seems to know what’s happened to them. (View Highlight)
  • Inspired by Dylan Freedman I tried the model using GanymedeNil/Qwen2-VL-7B on Hugging Face Spaces, and found that it was exceptionally good at extracting text from unruly handwriting: (View Highlight)
  • The model apparently runs great on NVIDIA GPUs, and very slowly using the MPS PyTorch backend on Apple Silicon. Qwen previously released MLX builds of their non-vision Qwen2 models, so hopefully there will be an Apple Silicon optimized MLX model for Qwen2-VL soon as well. (View Highlight)