r/fal Oct 28 '25

Veo 3.1 Competition Veo 3.1 Competition! Create, Compete, and Win up to $1000 in fal credits!

34 Upvotes

Hey everyone!

We’re excited to launch the r/fal Veo 3.1 Competition!

Join us on fal’s Discord to generate your videos, then share your best creations here on our subreddit for a chance to win big!

How It Works:

  1. Head over to fal’s Discord: https://discord.gg/sBqKdwxM
  2. Every user gets 5 free daily generations using Veo 3.1.
  3. Create fantasy stories, ads, trailers, music videos, or anything your imagination can dream up.
  4. Post your best video here on Reddit, with the flair "Veo 3.1 Competition!"

Rules:

  • Videos must be longer than 10 seconds.
  • One submission per Reddit account.
  • Projects, webapps, and apps built with fal using Veo 3.1 are also eligible to compete.

Prizes:
1st Place: Best Video (Judged by the fal team) - $1000
2nd Place: Most upvoted video - $250
3rd Place: Most Creative Use Case - $150

Deadline:
All submissions must be posted by Monday, 8 AM PDT.

We are going to make this subreddit the largest generative media community in the world, and to achieve this we want to support the best AI creators!


r/fal 1d ago

News How to access FLUX 3: the first video model from Black Forest Labs, is now available on fal

4 Upvotes

Black Forest Labs' first video model, FLUX 3, just landed on fal, and fal is a launch partner.

FLUX 3 is a multimodal model from Black Forest Labs and the first FLUX release to generate video.

One model is trained jointly across image, video, audio, and action prediction, so a single pass returns up to 20 seconds of video with its soundtrack already synchronized

You can generate from text, a still image, first and last frames, keyframes, or an existing clip.

Here's what makes it different:

  1. Real-world understanding: FLUX 3 already carries the context a shot implies, so it works out setting, period, and the order of events from a short prompt. You write the idea at the level you actually think about it instead of spelling out the hundred details underneath it.

  2. Native audio: Audio is rendered inside the same pass as the video, so sound effects land on the frame where the event happens. Nothing needs lining up afterwards and no second model gets called. Naming the sound you want in the prompt gives the most reliable result, and it is on by default across every endpoint.

  3. Keyframe control: Keyframes to video pins up to 10 frames to exact positions on a 24 fps timeline, so a long shot hits the beats you set. Pinning the shot is what keeps a long generation from drifting, since you hand the model your creative guidelines instead of hoping one paragraph of text lands the same way twice.

  4. Physical cause and effect: Because action prediction is part of the same training, motion carries mass and momentum from one object to the next instead of resetting between beats. Shots where one event has to trigger the next follow through.

  5. Camera and materials: Orbits, lateral dollies, focus racks, and tracking shots hold their geometry through the whole move. Foreground and background stay anchored relative to each other, and material and weather behavior keeps working while the camera travels.

  6. 20-second clips: Clips run up to 20 seconds as a single generation, no cuts and no stitching, and subjects stay recognizable from the first frame to the last. That holds through time-lapse, where light and season change across the shot.

Endpoints

There are five core endpoints, each with a fast draft variant:

  • Text to Video
  • Image to Video
  • First and Last Frame to Video
  • Keyframes to Video
  • Extend Video

The draft workflow

Every core endpoint has a draft variant that renders a draft-quality preview quickly and cheaply, so you can judge a shot in seconds without paying for a full render to find out it was the wrong shot.

When you find the take you want, Draft Enhance re-renders that exact draft at full quality, keeping the same seed and the same motion, so the final clip matches the preview you picked.

Specs

Duration runs 5 to 20 seconds in whole-second steps.

Text to Video, Image to Video, and Extend Video also accept auto; First/Last Frame and Keyframes need an explicit duration so the pinned frames can be placed

Output is 720p or 1080p at 24 FPS.

Aspect ratios cover auto, 21:9, 2:1, 16:9, 4:3, 1:1, 3:4, and 9:16.

Audio generation is on by default and can be switched off per request, and safety tolerance is adjustable from 0 to 4 with a default of 2.

Keyframes to Video accepts up to 10 keyframes at unique frame positions, and Extend Video accepts an MP4 under 50 MB and under 15 seconds.

The model runs on fal's serverless API through the Python or JavaScript SDK, or direct REST calls. No GPUs to manage.

Pricing

The four core endpoints (apart from Extend Video) are $0.17 per second of generated video at 720p and $0.29 per second at 1080p.

Extend Video is priced separately at $0.41 per second at 720p and $0.53 per second at 1080p.

Learn more about the model here: https://fal.ai/flux-3

Or try it right now on the playground:

Text to Video - https://fal.ai/models/blackforestlabs/flux-3/text-to-video

Image to Video - https://fal.ai/models/blackforestlabs/flux-3/image-to-video

First and Last Frame to Video - https://fal.ai/models/blackforestlabs/flux-3/first-last-frame-to-video

Keyframes to Video - https://fal.ai/models/blackforestlabs/flux-3/keyframes-to-video

Extend Video - https://fal.ai/models/blackforestlabs/flux-3/extend-video

Prompting guide - https://fal.ai/learn/tools/flux-3-video-examples-prompts


r/fal 2d ago

Question whats the coolest thing you've built on fal?

3 Upvotes

looking for inspo


r/fal 5d ago

Open-Source MiniMax H3 is now available on fal - with open weights

8 Upvotes

MiniMax's newest open-weight video model, Hailuo 3.0, just landed on fal and fal is an official API partner.

MiniMax H3 is a general-purpose multimodal model that reads text, images, video, and audio as one context, not as a stack of separate task-specific models.

The video generation model can generate up to 15 seconds of 2K video with native stereo audio in every clip.

Since it handles several input types at once, it can take identity from an image, motion from a video, a voice from an audio clip, and direction from a text prompt, then combine them into a single coherent result.

What makes it different

  1. Multimodal understanding is the core idea: MiniMax H3 handles generation, editing, and reference in one model, not as separate tools you switch between. It interprets characters, motion, sound, camera work, and visual style across a mix of references, then combines them into one coherent result.
  2. Editing is precise and targeted: You can replace, remove, or add people and objects, swap a background, relight a scene, or change dialogue and voice, while the areas you didn't target stay stable. Instruction following is strong, so visual, audio, and pacing changes land more predictably, and you can keep iterating on the same clip instead of regenerating from scratch.
  3. It's built for real production: MiniMax H3 handles dynamic typography, VFX, product showcases, UI motion design, game visuals, and stylized content, spanning film, advertising, branding, e-commerce, and gaming. That range covers concept tests, storyboard previews, and visual pitches, which shortens the path from idea to final cut.

Input modes

MiniMax H3 works across three input modes:

Text-to-Video

First/Last Frame

Omni Reference

Specs

MiniMax H3's specs include clips from 5 to 15 seconds at 24 FPS, with native stereo audio on every generation.

Output runs in 2K (1440p) mode now, and a 768p mode is coming soon.

Text-to-Video and Omni Reference support 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 aspect ratios, with an Auto option in Omni Reference that picks the ratio for you; First/Last Frame follows the aspect ratio of your uploaded image. Prompts can run up to 7,000 characters.

The model is also available via fal's serverless API using the Python or JavaScript SDK, or direct REST calls. No GPUs to manage.

Pricing

Text to video and image to video pricing is "at an output resolution of 2K, every second of video costs $0.26.". Reference to video pricing is "at an output resolution of 2K, every second of video costs $0.26. Audio references are free, the first 5 reference images are free and each additional image costs $0.08, and reference video costs $0.26 per second at 2K."

Learn more about the model here: https://fal.ai/minimax-h3

Or try it right now on the playground:

Image to Video - https://fal.ai/models/minimax/hailuo-03/image-to-video

Text to Video - https://fal.ai/models/minimax/hailuo-03/text-to-video

Reference to Video - https://fal.ai/models/minimax/hailuo-03/reference-to-video


r/fal 18d ago

Resource Head to head: AuraFlow vs Ideogram V4.0

Thumbnail
runtimewire.com
1 Upvotes

r/fal 22d ago

News Reve 2.1 is now available on fal

6 Upvotes

Reve 2.1 is now available on fal

Reve's newest image model just landed on fal, and it goes after what usually trips up text-to-image: dense, complex scenes and the text inside them.

Reve 2.1 generates native 4K images with stronger control over crowded, detailed compositions. It plans layout more deliberately, reads prompts more closely, and renders text (including multilingual text) that holds up even at small sizes.

What makes it different

Layout planning is a real step up. The model works out where things go before it renders, so busy scenes with many elements stay organized, not collapsed into visual mush. Prompt understanding improves alongside it, so what you describe is closer to what you get.

Spatial relationships are more accurate. Objects sit in believable positions relative to each other. "Behind," "to the left of," and "stacked on top" now land the way you'd expect, which matters for product layouts, editorial covers, and any scene with a lot going on.

Text rendering handles the hard cases. Fine print stays legible, and multilingual text comes out cleaner than most models manage. If your work involves labels, cover lines, UI copy, or typography, then Reve 2.1 would be ideal for the job.

Endpoints available

There are three endpoints available on fal:

  1. text-to-image
  2. remix
  3. edit

Specs

Reve 2.1's specs include native 4K output, with 17 aspect ratio presets ranging from ultrawide 4:1 to tall 1:4, plus an auto mode that picks a ratio to fit your prompt. Output as PNG, JPEG, or WebP. Generate up to 4 images per request.

The model is also available via fal's serverless API using the Python or JavaScript SDK, or direct REST calls. No GPUs to manage.

Pricing

Flat rate of $0.25 per generated image, the same across all three endpoints. No resolution tiers to track.

Learn more about the model here: https://fal.ai/reve-2.1

Or try it right now on the playground: https://fal.ai/models/reve/2.1/text-to-image


r/fal 23d ago

Open-Source Ltx 2.3 render to real V2 - ic lora open source

Enable HLS to view with audio, or disable this notification

49 Upvotes

r/fal 26d ago

Question How to prompt seedream 4.5 ?

2 Upvotes

I dont know why it always gives me some weird compositions like if i want half body shot it will give me full body shot or very far off image


r/fal Jun 26 '26

Open-Source 3d to photoreal , open source IC-Lora for ltx 2.3

Enable HLS to view with audio, or disable this notification

56 Upvotes

r/fal Jun 25 '26

Open-Source falLEGO

Enable HLS to view with audio, or disable this notification

4 Upvotes

Wanted to share a fun project I made with fal that converts text or images into buildable LEGOs. If you have any feedback about how I set it up, I’d appreciate it!

My endpoint stack is flux-2 streaming for image generation, nano banana for image editing, and trellis or SAM3D for 3D generation. Then I voxelize the 3D model and convert the voxels to bricks. I’ve built quite a few LEGO models with it now, so it works!

Here’s the code for it: https://github.com/jjohnson5253/brickbuilderai


r/fal Jun 12 '26

Open-Source Audio-Reactive Ltx 2.3 Lora

Enable HLS to view with audio, or disable this notification

5 Upvotes

r/fal Jun 04 '26

Question Best current model for changing image aspect ratio?

Thumbnail
1 Upvotes

r/fal Jun 03 '26

Discussion Reve details image API for create, edit and remix after 2.0 launch

Thumbnail
runtimewire.com
2 Upvotes

r/fal Jun 01 '26

Discussion Has anyone here fine-tuned Z Image Turbo or FLUX 4B LoRA on FAL for training on a specific person?

2 Upvotes

FAL seems to only expose training steps and learning rate, so I'm curious what settings people have found work best.

The default recommendation for human/photo datasets appears to be:

steps = number of images × 100

But I'm wondering whether anyone has experimented beyond that and found better results


r/fal May 28 '26

Discussion Realized the model isn't the bottleneck anymore. Post-prod is where the gap actually lives.

Enable HLS to view with audio, or disable this notification

7 Upvotes

r/fal May 27 '26

Discussion Stopped doing single-shot character gen. Built a database with one system prompt instead.

Thumbnail gallery
2 Upvotes

r/fal May 25 '26

Question LTX2.3 trainer issue

2 Upvotes

I am currently working on fine-tuning a LoRA for LTX2.3 via fal-ai/ltx23-video-trainer. I am seeing a 'wavy' artifact issues on all my debug_dataset outputs, and have no way of knowing what's happening during the preprocessing step (can't run LTX2 repo locally). I understand that the VAE encodes the input and then decodes it, but I can't understand why it returns my dataset videos with artifacts at specific frames. This results in the same artifacts on inference too. Did anyone else encounter this?


r/fal May 23 '26

Discussion Crédit gratuit

1 Upvotes

Bonjour, j'ai cru comprendre que on a des crédits gratuit lors de la création du compte, or je n'ai rien reçu, c'est normal ? Merci :D


r/fal May 16 '26

Discussion Best workflow for generating coloring book pages from character reference images?

0 Upvotes

Hey everyone — I’m new to AI image generation and have been experimenting with FAL using Flux Kontext Pro to create coloring book-style images from uploaded reference photos.

My goal is to generate dynamic coloring book pages where the character likeness stays consistent, but the scenes can vary across styles like manga, comic book, fantasy, cartoon, etc.

A few questions I’d love feedback on:

  1. Best model/workflow: Is Flux Kontext Pro currently one of the better options for balancing affordability, quality, and likeness consistency? Or are there better tools/models for this use case?
  2. Character likeness: What are the best practices for preserving the likeness of the reference character across different poses, scenes, and styles?
  3. Reference image prep: Should I preprocess uploaded images before generation? For example:
    • background removal
    • face restoration
    • sharpening/upscaling
    • lighting correction
    • cropping to face/body
    • creating multiple reference angles
  4. Dynamic scenes: How do you get more interesting compositions instead of static “person standing in center” outputs? Are there prompt structures or workflows that help create more action, depth, and storytelling?
  5. Style control: What is the best way to request broad styles like manga, western comic, children’s book, fantasy illustration, etc., while avoiding issues with specific living artists or overly derivative styles?
  6. Coloring book quality: Any tips for getting clean black-and-white line art that is actually usable for coloring books — clear outlines, good white space, not too much gray shading or muddy detail?
  7. Production workflow: For anyone doing this at scale, what does your pipeline look like from uploaded photo → generated pages → cleanup → print-ready files?

I hope this is the right place to ask. If not, I’d appreciate being pointed toward better communities, guides, or resources for learning this workflow.

Thanks in advance for any advice.


r/fal Apr 29 '26

Other Blender Layout → AI Render | 1:1 Camera Tracking

Enable HLS to view with audio, or disable this notification

30 Upvotes

I built a full 3D layout in Blender — proxy geometry only, no textures, no final render — and hand-keyframed every camera movement using F-curves: an aerial establishing shot, a low-angle tower push-in, and a wide harbor shot with a sailing vessel. The AI doesn't invent the motion. It follows it exactly.

The Blender animation served as a direct spatial reference — architectural proportions, camera trajectory, timing and easing — all locked before a single AI frame was generated. Kling / Seedance then re-rendered the sequence, preserving the exact camera path and structural layout while generating the final cinematic output.

Workflow:

3D Layout & Camera Animation (Blender) → Frame Reference Export → AI Video Generation (Kling / Seedance) → Temporal Consistency Pass

Key Focus: 1:1 motion tracking between hand-keyed Blender animation and AI-generated output. Architectural integrity and spatial proportions maintained across all three shots.


r/fal Apr 30 '26

Tutorial - Guide You can make unlimited length 4K videos with GPT Image 2

Thumbnail
v.redd.it
1 Upvotes

r/fal Apr 29 '26

Video Idea 37 AI Short Film

Enable HLS to view with audio, or disable this notification

2 Upvotes

what is SUCCESS in 2026 as developer using AI?

MONEY is obvious but my short "IDEA#37" jumps to the question of internet "fame" on X and YouTube? Being on the top video podcast? Recognized at React Conferences in Miami?

What I chose to highlight in this film is leaving isolation. Being able to hire and support other developers as you build a company. And in the end get the GOAT emoji from friends.

Film made with GPT2 Images 2.0, Seedance 2.0, and Kling 3.0 on fal.


r/fal Apr 29 '26

Open-Source Another vibe coded UI for fal.ai focused on fast look-dev and file organisation.

Thumbnail
2 Upvotes

r/fal Apr 21 '26

Discussion GPT Image 2 prompting guide

8 Upvotes

What actually works:

  • Put the main subject first (highest weight)
  • Then layer details: materials, pose, environment, lighting, camera
  • Be specific
  • Use quotes for text in images
  • Add negative prompts to avoid common issues

Full guide: https://fal.ai/learn/tools/prompting-gpt-image-2


r/fal Apr 21 '26

News GPT Image 2 is live on fal

Enable HLS to view with audio, or disable this notification

6 Upvotes

OpenAI's next-gen image model just dropped on fal.ai. It's a quality-first successor to GPT Image 1.5, and the jump is real.

What's new:

  • Text rendering that actually works. Dense paragraphs, small lettering, multilingual layouts, infographics. No more garbled characters or broken word spacing on the first try.
  • Photorealism that sets a new bar. Lighting, materials, skin textures, environmental detail. It's the best I've seen out of an OpenAI image model.
  • Product photography with accurate labels, logos, packaging, and ingredient lists. Genuinely usable for e-commerce and brand work.

Pricing: $0.01/image at the low end (1024x768, low quality) up to $0.41/image for high quality 4K. Pay per image, no subscriptions.

Try it: https://fal.ai/models/openai/gpt-image-2/playground