Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
laion/clap-htsat-fused · Discover · Kaino
Discover/MODELS/laion/clap-htsat-fused
laion/clap-htsat-fused logo

MODELS

laion/clap-htsat-fused

by LAION

modellead-sourcehugging-face-popular-modelssource:github.comaudioaudio-classificationmultimodallanguage-audioclaphugging-facetransformerslaion
Visit WebsiteDocumentationGitHub

Overview

LAION CLAP model checkpoint for multimodal audio-and-language representation and audio-classification workflows.

Details

laion/clap-htsat-fused is a LAION model checkpoint listed on Hugging Face for audio-classification. Hugging Face Transformers documents CLAP as a multimodal audio-and-language model and includes examples that instantiate ClapAudioModel and ClapProcessor from laion/clap-htsat-fused. The official LAION-AI CLAP repository describes the project as providing audio and text representations via Contrastive Language-Audio Pretraining and links the implementation to LAION and the CLAP paper.

When to Use

Use when you need a CLAP checkpoint from LAION for audio-and-language representation tasks. Use with Hugging Face Transformers examples that load ClapAudioModel and ClapProcessor from laion/clap-htsat-fused. Evaluate for audio-classification pipelines where a Hugging Face-hosted CLAP model checkpoint is appropriate.

Getting Started

  1. Open the Hugging Face model page at https://huggingface.co/laion/clap-htsat-fused to review the model card and files.
  2. Read the Hugging Face Transformers CLAP documentation for usage patterns with ClapAudioModel and ClapProcessor.
  3. Review the LAION-AI/CLAP GitHub repository for project context
  4. implementation details
  5. and links to the CLAP paper.
  6. Run a small audio-and-language or audio-classification evaluation before relying on the checkpoint in a production workflow.

Key Features

  • •Hugging Face model checkpoint published under laion/clap-htsat-fused.
  • •Associated with CLAP
  • •described by Hugging Face Transformers as a multimodal audio-and-language model.
  • •Supported in Transformers documentation through examples using ClapAudioModel and ClapProcessor.
  • •Connected to LAION-AI's CLAP repository for Contrastive Language-Audio Pretraining.

Capabilities

  • •audio-classification
  • •audio-and-language modeling
  • •audio representation
  • •text representation
  • •contrastive language-audio pretraining

Last updated Jun 4, 2026