Two cats and two dogs
The four benchmark subjects varied species, fur, pose, and setting so the launch choice was not based on one flattering pet image.
AI and quality methodology
The current launch system uses a pinned DreamActor M2.0 model through Replicate for motion transfer. PetShimmy controls the movement, validates the file, adds the exact matching audio, and automatically delivers the private result after its technical checks pass.
Method reviewed 31 July 2026
The production path
Each boundary has a different job: reject unsuitable input, control the motion, verify the provider output, assemble the audio, validate the final file, and decide whether a customer should receive it.
PetShimmy accepts one portrait JPEG, PNG, or WebP. It checks the real file structure and dimensions, then prepares a normalized pet image for the model.
The chosen dance and purchased duration must match a registered motion asset. Its location, type, size, dimensions, timing, and SHA-256 are checked before use.
Replicate receives the normalized pet image and the verified silent reference for the pinned DreamActor M2.0 version. PetShimmy does not rely on a free-form text prompt to invent the choreography.
The returned media must come from a trusted provider host and pass byte-size, MP4, duration, portrait, codec, and minimum-resolution checks. Any provider-returned audio is discarded.
PetShimmy combines the video with the exact music or voiceover companion prepared for that performance and duration, then verifies video, audio, duration, track continuity, and timeline alignment.
After the final file passes integrity, duration, portrait-format, video, and audio checks, PetShimmy marks it ready automatically and emails the customer’s secure result link.
The launch benchmark
The retained launch benchmark used five dances, four pet subjects, two durations, and two independent seeds: 5 × 4 × 2 × 2 = 80 DreamActor M2.0 outputs.
The four benchmark subjects varied species, fur, pose, and setting so the launch choice was not based on one flattering pet image.
All 80 recorded calls returned an output and every output was manually reviewed. Completion is reliability evidence—not a claim that every clip was fit to deliver.
The five music-led dances progressed from this benchmark with mandatory finished-output review. Full Rant was assessed separately and showed greater anatomy variation on dogs, so its outputs use the same review and refund boundary.
Why publish the caveat? A provider success rate and a customer-usable delivery rate are not the same measure. PetShimmy does not turn “80 jobs completed” into “80 perfect videos.”
Known generative limits
The model interprets human movement for an animal body. Timing, poses, expression, paws, limbs, fur, anatomy, and background details can change. Small variation is expected; a substantial defect is a quality issue.
The audio does not come from the model. PetShimmy uses the prepared music or voiceover for the selected performance and duration and verifies it against the video timeline.
Automated checks can prove file integrity, duration, and tracks. A person decides whether the pet remains broadly recognizable and the dance is clear enough to deliver.
Judge the output yourself
Watch five actual six-second dance results and the complete 16-second Full Rant, with matching audio, before choosing.