jdopensource/JoyAI-Echo
text-to-video · ltx-video, JoyAI-Echo, text-to-video
Home / 🎬 Media Creation
60 View all · Daily curated AI / LLM open-source intelligence, with plain-language notes and license checks.
text-to-video · ltx-video, JoyAI-Echo, text-to-video
text-to-audio · stable-audio-3, safetensors, audio-generation
text-to-speech · safetensors, moss_tts_delay, text-to-speech
text-to-speech · transformers, safetensors, qwen3_5_text
text-to-image · diffusers, text-to-image, lora
text-to-image · diffusers, krea2, text-to-image
image-to-video · diffusers, safetensors, i2v
text-to-speech · supertonic, onnx, text-to-speech
text-to-audio · magenta-realtime-2, tflite, text-to-audio
image-to-video · cosmos, diffusers, safetensors
image-to-image · controlnet, lora, depth
automatic-speech-recognition · transformers, safetensors, cohere_asr
image-to-video · lora, ic-lora, ltx
image-to-image · Flux2Klein, Sun, I2I
text-to-speech · transformers, safetensors, higgs_multimodal_qwen3
text-to-image · cosmos, diffusers, safetensors
image-to-video · diffusers, safetensors, i2v
text-to-image · diffusers, text-to-image, lora
text-to-video · ltx-video, identity-preservation, ipt2v
text-to-image · diffusers, safetensors, text-to-image
text-to-image · diffusers, text-to-image, lora
text-to-image · diffusers, safetensors, text-to-image
text-to-speech · pytorch, text-to-speech, tts
image-to-video · diffusers, character-animation, video-generation
text-to-speech · pytorch, safetensors, text-to-speech
text-to-image · diffusers, safetensors, ternary
text-to-video · custom, ti2v, text-to-video
text-to-video · diffusers, safetensors, gguf
text-to-speech · ZONOS2, text-to-speech, license:apache-2.0
image-to-image · pytorch, diffusers, safetensors
automatic-speech-recognition · nemo, safetensors, nemotron3_5_asr
automatic-speech-recognition · pyannote-audio, pyannote, pyannote-audio-pipeline
text-to-image · diffusers, safetensors, text-to-image
text-to-speech · safetensors, qwen3_tts, text-to-speech
AI Image Generator is a powerful desktop application for creating stunning AI-generated artwork from text prompts. It integrates with Stable Diffusion XL, Flux, DALL-E 3, and Midjourney APIs for both local GPU-accelerated and cloud-based image generation.
Professional color grading, video editing, and visual effects software featuring advanced HDR tools, real-time AI tracking, and hardware-accelerated processing pipelines.
乔木智能视频导演 Skill:素材治理、双语字幕、品牌包装与可复现渲染 | Agent-native video director with governed sourcing and verifiable rendering
Local-first conversational AI video editor with a professional multi-track timeline, Agent Skills, MCP integration, and Remotion-powered rendering.
半调纸拼贴 B-roll 生成 skill:三闸门审批,Gemini Omni Flash 首尾帧组装动画 | Editorial halftone paper-collage B-roll agent skill
Stable Diffusion WebUI Portable with model pack, ControlNet, LoRA library, and extensions—full local AI art studio unlocked.
Sony Vegas Pro 21免費版的影片編輯和後製工具。
Turn one topic into a finished Vox-style paper-collage explainer/ad video — automated end to end on Atlas Cloud + ffmpeg. An agent skill.
Stable Diffusion Free Local - run Stable Diffusion locally for free AI image generation.
Clone a voice from a few minutes of audio and generate speech locally — Qwen3-TTS fine-tuning pipeline with CLI and web UI
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
stable diffusion web ui free comfyui sd webui video git clone github ai art generation image generator local deployment python 3.10 gradio interface txt2img img2img inpainting outpainting controlnet ip-adapter lora checkpoint model loader embeddings textual inversion lycoris civitai hugging face extension manager
AI-powered image generation toolkit using Stable Diffusion models for creating stunning visuals from text prompts
Agent-native AI video factory: one brief → narrated, subtitled, scene-editable videos (shorts & long-form). Deterministic HTML rendering, free keyless stack, Korean-first.
Zero-dependency browser video editor that AI agents can drive — JSON timeline, MCP + REST, live-reloading UI
Unified World Model Inference & Evaluation Infrastructure
Opensource, on-device voice + vision layer for AI agents. Bring any model or coding agent; the whole speech loop (VAD, STT, TTS, barge-in) runs locally. An open alternative to ElevenLabs Agents, Gemini Live, and OpenAI Realtime.
audio.cpp with a full-task WebUI - pure C++ audio-model inference engine powered by ggml. TTS, ASR/STT, VAD, voice conversion, speaker diarization, music generation. No Python dependency.
Pinnacle Studio 相關的音訊編輯指南與顯示設定。
未提供描述。
DaVinci Resolve Studio workflow — color grading, Fusion comps and delivery presets on Windows.
Free open-source project designed for turning youtube-viedos into viral short videos. Highlight detection, subtitles, translation, voiceover, all in one for your content.
Motion graphics automation framework for Adobe After Effects CC 2024+. Procedural keyframe generation, expression-based animation control, layer automation, and batch rendering. Scripting tools for animators and motion designers. Includes keyframe generators, expression builders, and render automation.
Research-backed agent skills and tools for premium image, video, audio, voice, and generative media production across AI coding assistants.
Turn your coding agent into a video studio: describe a video in plain language, and your agent writes the timeline and produces the file.
Talk to 峰哥 — 克隆任何人的声音和性格,实时语音对话,工程延迟 < 1 秒 | Clone anyone's voice & personality for real-time conversation. < 1s engineering latency.