Hugging Face released Transformers v5.14.0, adding the Inkling multimodal model from Thinking Machines (975B total, 41B active parameters) and significant Flash Attention performance improvements up to 260% for large inputs, boosting enterprise AI capabilities in MENA.

1 min read

Hugging Face Transformers v5.14.0 Released with Inkling Multimodal Model (975B Parameters)

Release Overview

FAQ

What is Hugging Face Transformers v5.14.0?

It is a new release of the Transformers library adding the Inkling multimodal model from Thinking Machines (975B total, 41B active) and major Flash Attention performance improvements and MTP support.

What is the Inkling model?

A general-purpose multimodal model from Thinking Machines with 975B total parameters (41B active) accepting text, image, and audio inputs, suitable for enterprise applications in the region.

How does this release affect AI teams in MENA?

It provides up to 260% performance improvements and powerful multimodal models, reducing infrastructure costs and accelerating deployment in enterprise and government applications.

What are the breaking changes in this release?

GPTNeoX now remaps embed_out to lm_head, and GPTBigCode has _supports_attention_backend = True enabled for vLLM compatibility, requiring code updates for existing users.

Source: Hugging Face Transformers Releases

AI-assisted content, human-reviewed.