Auto Secure LoginPlatform
Early accessWebWindowsAI you can own

Clone your voice without giving it away.

One audio workspace for voice, speech, transcripts, music, mixing and evidence, with no third-party AI service involved.

Built for: Voice actors and narrators who want a usable copy of their own voice without signing it over to somebody else · Podcasters and audiobook producers who need to fix three words without rebooking a session · Small video, animation and game teams building scenes where several characters talk to each other · Musicians and producers who want a sung guide vocal or an instrumental sketch before they book players

17.95 sFirst music render after a cold restart, measured end to end over the live site
6.238 sRepeat three-second music render, measured, producing valid audio both times
20.45%Music render time cut by a measured settings change, with identical audio bytes before and after
30.212 sA five-second music render including a cold model load, measured on the machine that runs the engine

The problem

ASL Voice Lab exists because of this.

You needed one line fixed. The narrator has moved on, the studio is booked out three weeks, and the client wants it Thursday. So you look at voice cloning, and every option asks you for the same thing first: upload a clean recording of the voice, agree to terms you did not write, and trust that the recording stays where they said it would. There is usually no delete button that deletes anything. There is rarely a straight answer about whether your voice becomes training data. And the voice you get back is theirs to switch off, reprice, or retire whenever they choose. For a working narrator, that is not a tool. It is a lien on the only asset they have. Meanwhile the actual job is spread across five subscriptions. One place makes the speech. A second one transcribes. A third writes captions. A fourth is the mixer. A fifth generates a bed of music. Every handoff is a download, a re-upload, a filename you will not recognize next week, and another copy of confidential audio sitting on another company's disk. By the time the piece is finished you could not honestly say where all the copies are, which version was the final one, or what made any given clip. Then there is the other half of the problem, the one nobody sells you a clean answer to. Somebody sends you a recording and asks whether it is real. The tools that claim to answer that hand back a confident number - eighty-seven percent likely synthetic - produced by a model that has never seen the generator in question, on audio a messaging app has already re-encoded twice. That number will not survive a cross-examination and it should not survive yours. What you actually need is a record of what can be measured, an explicit statement of what cannot, and a file you can hand to somebody else.

What changes

  • Your reference recording is encrypted at rest with AES-256-GCM, and revoking a voice deletes the recording from disk rather than flagging it
  • No outside voice, transcription or music company is ever called - if a required model is missing the request stops with a clear error instead of quietly routing your voice somewhere else
  • Everything it generates carries a signed receipt inside the audio file, so a clip can prove where it came from later
  • The analysis room returns inconclusive and explains why, instead of a confident percentage that would not survive scrutiny
  • Eight rooms that hand audio and text to each other directly - no download-and-re-upload between transcript, script, mixer and analysis
  • Repeatable by design: the same voice, text and seed produced byte-identical files with matching hashes in the recorded acceptance runs, which is what makes re-rendering an approved script safe
  • Music that is not cleared for commercial release is labeled that way on the screen, in the response and inside the receipt, not buried in a terms page

What it does

How ASL Voice Lab works, start to finish.

ASL Voice Lab is one workspace with eight rooms - Voices, Speak, Perform, Script, Transcribe, Music, Mix and Analyze - and the rooms hand work to each other without a single download-and-re-upload in between. You record one clean voice sample, about ten seconds. That sample is encrypted the moment it is stored and becomes a voice profile you can use, list, and revoke. Revoking is not a flag in a table: it deletes the stored reference recording from disk. Nothing about that profile is pooled, shared, resold, or used to improve anything for anybody else, and no outside voice, transcription or music company is ever called. With a voice active, Speak turns typed text into audio in that voice. Long scripts are divided at natural sentence breaks and each section starts playing as soon as it is ready, so a three-paragraph passage is audible while the rest is still rendering. You are not watching a progress bar for a minute only to discover you typed the wrong name. Voice Director gives you the controls a director actually uses: pace, the gap between sections, a delivery direction, a pronunciation dictionary for the words your product name breaks, and the ability to re-render one selected phrase instead of the whole passage. Performance Mode takes a different route: you act the line yourself, the workspace writes down the words you said, re-speaks them in the voice profile you chose, and then stretches that render onto the exact duration and loudness shape of your read. The delivery is yours even when the voice is not - and the product is careful to say that alignment is not proof of identity. Script builder handles up to twelve lines with a different voice on each, a pause of nought to ten seconds after any line, and one starting seed that gives every line its own repeatable seed, then places the finished scene on the mixer as separate signed tracks without disturbing anything already there. Transcribe works both directions. Record straight from a microphone you pick by name, with a live input level so you can see you are being heard, or drop in a file in any of nine common formats. You get the text, the detected language, the duration, and word-level timings you can click to seek the audio, with downloads as plain text, JSON, SRT or WebVTT captions. You can edit the transcript in place, and those edits stay in your browser - they are never sent back. Live captions run segmented while you talk, with speaker labels you type yourself. Audio and transcript text are not retained by default, and the temporary working copy is deleted as soon as the job returns. When you have generated speech from a script, one button compares your source text against an independent transcript of the result and gives you a downloadable word-agreement report - a real check on whether the render actually said what you wrote. Mix is a twenty-four track mixer that runs in your own browser, laid out against a tempo, key, meter and bar grid it can detect for you. Trim, fade, loop, solo, mute, gain, pan, timeline offset, split markers, crossfades, playback rate and pitch, EQ, compression, de-essing, reverb, delay, waveform views, undo, and frequency-band stems. All of it non-destructive, none of it touching your source files, none of it uploaded to work on. The Music room adds two things: a sung vocal built from one of your own voice profiles against a typed note pattern or a standard MIDI melody, and an instrumental generator that will create new material, continue what you have, replace a section, or build a clean-repeating loop, with up to four takes from one prompt. Only when you export does a finished mix leave your device, and only so it can be signed. The whole project - settings, tracks and results - autosaves locally and exports as one encrypted file under a passphrase you choose. Analyze is the room that refuses to guess. Send it any audio and it checks for a valid signed receipt, looks for a supported watermark, and maps signal measurements over time: anomaly scores per window, spectral discontinuities, and optional acoustic similarity against one of your saved voices. What it will not do is turn that into a verdict. It withholds attribution of the source generator. It makes no speaker-identity call. When there is no provenance to verify it returns inconclusive and says so in plain words on the screen, including the line that matters most - no watermark found never means real. The whole evidence set, uncertainty statement included, downloads as a JSON file that does not carry the audio inside it. Every piece of generated speech, every performance transfer, every sung vocal, every generated instrumental and every exported mix carries a signed receipt embedded in the audio file itself, so a clip that came out of this workspace can prove it later. The product is equally clear that a receipt can be stripped by re-encoding, which is exactly why its absence proves nothing.

Features

Everything in the current release.

Each of these is built and working today. Nothing on this list is a roadmap item.

01

One recording, one voice, ten seconds

Name the voice, say whether it is yours or another speaker's, and speak for six to twelve seconds. It keeps the clearest ten-second stretch, mixes it to mono, resamples it, and encrypts it before it is stored. If the sample is mostly silence it tells you rather than building a poor voice from it. One good take is the whole enrollment - there is no back catalogue to upload.

02

Revoking a voice deletes the recording

Profiles are listed as active or revoked, and revoking one removes the stored reference recording from disk rather than hiding it behind a status flag. The button says so before you press it. This is the part most voice tools stay quiet about: if you decide next month that the voice should not exist here, it does not.

03

Long scripts you can hear while they render

Paste up to 2,500 characters. It divides the text at natural sentence boundaries, and each finished section starts playing while the later ones are still being made. You see which section is running, the elapsed time, an estimate to finish, and your place in the queue. You can stop after the current section and keep everything already done, and the completed sections are joined into one signed file.

04

Voice Director, for delivery rather than words

Set the pace, the gap between sections, and a direction: natural, warm, energetic, serious or intimate. A pronunciation dictionary fixes the words that always come out wrong, one written-equals-spoken rule per line, saved with the project. If a single phrase is off, select it and re-render that phrase instead of the whole passage.

05

Performance Mode keeps your read

Record yourself acting the line, or bring in a performance up to four minutes. The workspace transcribes what you said, re-speaks those words in the voice profile you chose, and aligns the result to your source duration and loudness shape. It reports how closely it aligned, and it deliberately stops short of claiming the result is emotionally or biometrically equivalent to you, because that would not be true. Because it works from a transcript, check the words before you use the take.

06

Twelve-line scenes, a different voice on each

Assign any active profile to each line, give it up to 1,000 characters and a pause of 0 to 10 seconds after it, and set one starting seed so every line renders the same way tomorrow. Lines are generated one at a time and land on the mixer as separate signed tracks. Anything already in the mix stays where it was; the scene is appended after it.

07

Transcription with word timings and four export formats

Record from a microphone you choose by name, or upload WAV, MP3, WebM/Opus, OGG, M4A/MP4, AAC, FLAC, 3GP or AMR up to fifteen minutes and 18 MB. You get text, the detected language, the duration, and word-level timestamps you can click to seek playback. Download it as TXT, JSON, SRT or WebVTT. Long leading and trailing silence is trimmed before recognition without moving your timestamps.

08

Your transcript edits stay in your browser

Fix the names and the jargon in place. Those corrections live on your device and are never sent back. TXT and JSON downloads carry your edits, while the caption files keep the original words and timings so subtitle timing stays trustworthy. Audio and transcript text are not retained by default, and the temporary working copy is deleted as soon as the job returns.

09

Live captions while you speak

Start captions and watch segmented text appear as you talk, with timestamps merged locally and speaker labels you type in yourself. It is honest about what that is: caption segmentation under your control, not automatic speaker detection, which is a far harder claim than most products admit to making.

10

A twenty-four track mixer in your own browser

Arrange voice, vocals, music and effects against a tempo, key, meter and bar grid, with a beat ruler and a starting tempo and key it can detect for you. Gain, pan, mute, solo, trim, fade, loop and timeline offset on every track. Nothing is uploaded to arrange it; the audio stays on your device until you export.

11

Non-destructive editing, effects and frequency stems

Split markers, crossfades, end automation, playback rate and pitch, three-band EQ, compression, de-essing, reverb, delay, waveform previews, undo, and frequency-band stems generated locally. Your source audio is never modified. Optional master compression and normalization are applied when the mix is rendered, not baked into the files you brought in.

12

Sung vocals and instrumental sketches

Turn an active voice profile into a note-timed sung vocal from a typed note pattern or a standard MIDI melody, with tempo, source pitch, vibrato and breath controls. The instrumental generator makes three to one hundred and twenty seconds of music - a three-second quick render is the default - and will create new material, continue what you have, replace a section, or build a clean-repeating loop, with up to four takes from one prompt.

13

An evidence workbench that returns inconclusive

Inspect any audio for a valid signed receipt, a supported watermark, time-window anomaly measurements, spectral discontinuities, and optional acoustic similarity to a saved voice. It withholds attribution of the source generator and makes no speaker-identity call. Where there is nothing to verify it says inconclusive, and it states on screen that no watermark found never means real.

14

Signed exports and portable encrypted projects

Generated speech, performance transfers, sung vocals, generated instrumentals and exported mixes all carry a signed receipt inside the audio file itself, naming what produced them. Your whole project - settings, tracks and results - autosaves locally and exports as a single encrypted file under a passphrase you choose, which you can archive and reopen later. Downloadable evidence JSON never embeds the audio it describes.

Proof

Numbers we can stand behind.

Every figure below comes from the product's own release record or test suite, not from a marketing estimate.

17.95 sFirst music render after a cold restart, measured end to end over the live site
6.238 sRepeat three-second music render, measured, producing valid audio both times
20.45%Music render time cut by a measured settings change, with identical audio bytes before and after
30.212 sA five-second music render including a cold model load, measured on the machine that runs the engine
342 msMedian transcription time for a 4.684-second reference clip across five runs
0.073Median real-time factor for transcription - roughly thirteen times faster than the recording is long
4,109 msMedian speech render in the measured winning thread configuration, chosen over four slower alternatives
25European languages the transcriber handles, with punctuation, capitalization and word timings
  • Mixing, editing, stems and mastering happen in your own browser; only the finished export is sent, and only to be signed
  • It can run entirely on a Windows machine you own, which is the version of this that keeps every recording inside your own network

Where it runs

Surfaces and status.

WEB
Web · 1.3.0 in the repository; the deployed build's version cannot be read from outside the sign-in gateEarly access - the full workspace is deployed and answering over HTTPS behind a single shared sign-in, deliberately kept out of search results; there is no self-serve sign-up and no individual accountsvoice.autosecurelogin.com
WIN
Windows · 1.3.0Self-hosted - the workspace runs on a Windows machine you own from the included launch script, which is the setup that keeps every recording inside your own network; you install and maintain the speech, transcription and music engines yourself, and there is no packaged one-click installer or download page

Status as of 2026-09-02. Checked live on 2026-09-02: voice.autosecurelogin.com is deployed and answering over HTTPS, returning an authentication challenge rather than an error, with headers that keep it out of search results. The repository version reads 1.3.0 and its most recent change is dated 2026-08-30. Nine dated release records run from 2026-08-27 through 2026-08-30 covering transcription, multi-voice scripting, the mixer, sung vocals, local music generation, portable encrypted projects and deterministic multi-process planning, each with measured acceptance evidence recorded alongside it - including runs made against the live public site. It is not Live for two reasons the repository states itself: access sits behind one shared sign-in with no individual accounts, no self-serve sign-up and no published price, and the repository's own list of what a public release still needs includes per-user accounts, audit logging, rate controls, abuse review, an age and guardian policy, and stronger ownership verification. Nothing in the repository evidences a user outside the company, so Early access here means access can be arranged, not that customers are already working in it.

What is new

Recent progress.

This product ships often. The most recent verified changes, newest first.

  • and.

  • , each with measured acceptance evidence recorded beside it.

  • transcription arrived as an isolated local engine with pinned models, signed receipts and per-request deletion; the twelve-line multi-voice script builder landed with automatic signed placement on the mixer; a transcription performance release brought a controlled 4.684-second clip down to a 342 ms median at a 0.073 real-time factor, with a queue that admits normal bursts instead of rejecting them; and a voice performance release added sectioned long-form speech that plays while it renders.

  • a voice release settled on the measured four-thread configuration at a 4,109 ms median with in-memory caches, after six, eight, ten and twelve threads all measured slower; and version 1.0.0 added the workstation-only dataset and evaluation console, immutable rights-checked dataset snapshots, frozen evaluation plans and a measured hardware report on top of the song mixer, sung-vocal generator and local instrumental generation that had arrived in 0.9.0.

  • version 1.1.0 made a three-second quick render the default and cut render time by 20.45% with identical audio bytes; version 1.2.0 added encrypted portable projects with autosave, Voice Director, Performance Mode, segmented live captions, non-destructive editing with effects, automation, undo and local frequency stems, music create/continue/replace/loop with up to four takes, and expanded evidence measurements.

  • version 1.3.0 added deterministic multi-process planning for preprocessing and evaluation work. Two changes were tried and rejected in the same window, with the evidence kept rather than discarded - a narrower fallback that ran slower and altered the output bytes, and a startup warm-up whose eight-second cost was not justified by what users would actually feel.

Pricing

Pricing for ASL Voice Lab is quoted after a short conversation about your situation, because the right scope differs from one team to the next. There is no charge for that conversation.

Ask about pricing
Honest by default. We publish the standard each product meets and the limits of each safeguard next to the feature, not in a footnote. If you cannot find an answer on this page, the assistant in the corner reads only these pages and will say so rather than guess.

Questions buyers ask

Straight answers.

What does it cost?

No price is published yet. Access is arranged directly with us while we work with early users, and pricing will be set before it opens more widely. If cost is the deciding factor for you, say so when you get in touch - it is genuinely useful to know what a tool like this is worth against the stack of subscriptions it would replace.

We already pay for a cloud voice tool. Why would we move?

Two reasons, and neither is a feature list. First, in most cloud tools your reference recording is theirs to keep and the delete button is a status change; here it is encrypted at rest and revoking a voice removes the recording from disk. Second, this is one workspace instead of five subscriptions - speech, transcripts, captions, mixing, music and evidence hand off to each other directly, so confidential audio stops being scattered across other companies' disks.

Where does my voice recording actually go?

In the hosted workspace it goes to us and no further. It is encrypted the moment it is stored and only ever decrypted to reach the speech engine we run. It is not pooled, shared, sold, or used to train anything for anyone else, no outside voice company is ever called, and there is no quiet fallback: if the required model is not present, generation stops with a clear error. If it must not reach us either, the whole workspace can run on a Windows machine you own.

Can I use what I make commercially?

Speech generated from your own voice profile, your transcripts, and your mixes are yours. The instrumental music generator available today is the exception: its output is not cleared for commercial release, and that restriction appears on the screen, in the response, and inside the embedded receipt rather than hidden in a terms page. A commercially clear music model is being built from scratch on rights-checked material and stays disabled until both of its stages are trained and a frozen evaluation passes without regression.

Is it actually ready, or am I going to be the one finding the problems?

It is deployed and answering behind a shared sign-in, and every release since 2026-08-27 carries measured acceptance evidence - including runs made against the live site - alongside checked-in test suites that run before a release. What it is not yet is self-serve: there is no sign-up page, no individual accounts, no published price and no billing, and the build's own list of what a public release still needs includes audit logging and abuse review. If that matters to you more than early access does, wait for the open release and we will tell you when it lands.

Do I have to migrate my existing project into it?

No. Bring in individual files - nine common audio formats are accepted - and work on whatever you need. Your existing files are never modified, because editing is non-destructive and the originals stay exactly as they are. A whole project also exports as a single encrypted file you can archive and reopen, so nothing you make here gets trapped inside it either.

What happens to my voices and my work if we stop paying, or if you go away?

Your exports are ordinary audio files that play anywhere, your transcripts and captions are standard formats, and your project file is yours - keep the passphrase, because without it nobody can reopen it. Revoking a voice deletes the recording, so you can leave with nothing of yours left behind. The workspace can also run entirely on hardware you own, which is the version of this answer that does not require trusting us at all.

Can it tell me whether a recording is a deepfake?

It can tell you what is measurable and it will not pretend to more than that. It verifies a signed receipt when one is present, checks for a supported watermark, and maps anomaly scores, spectral discontinuities and optional acoustic similarity to a saved voice. It withholds attribution of the source generator, makes no speaker-identity call, and returns inconclusive when there is nothing to verify - because no watermark found never means real, and any tool handing you a confident percentage is selling you something that will not hold up.

Can we run it inside our own building?

Yes, and that is the setup we would point you to first if your recordings genuinely cannot leave your network. The workspace and the speech, transcription and music engines all run on a Windows machine you own, reachable only from that machine. It is a real installation rather than a one-click setup - you install and maintain the engines yourself - so plan for that, not for an afternoon.

What should I know before I rely on it?

We would rather you hear this from us than discover it later. As of 2026-09-02:

  • Access is arranged directly with us and sits behind a single shared sign-in. There is no self-serve sign-up, no individual user accounts, and no published price.
  • In the hosted workspace your audio does leave your device. Enrollment recordings, transcription audio and anything you send to the analysis room travel to engines we run - not to any outside AI company. Mixing and editing are the exception: they happen in your browser, and only the finished export is sent, to be signed. If nothing may leave your network at all, run the workspace on your own hardware.
  • Speech, transcription and music are produced by engines on a machine we operate. When it is offline those rooms are unavailable and the interface says so, rather than falling back to an outside service.
  • This is not an identity or authentication system. A voice profile is a creative asset, not proof of who somebody is, and the product says so on the screen.
  • Enrollment asks whether the voice is yours or another speaker's, but it does not verify ownership. Stronger ownership verification, per-user accounts, audit logging, rate controls, abuse review and an age and guardian policy are all on the repository's own list of what a public release still needs.
  • The analysis room is not a deepfake detector. It gathers evidence and states uncertainty. No classifier is switched on here, and none will be until it has been measured across codecs, speakers, languages, noise, replay, editing and generators it has never seen.
  • Absence of a signed receipt or watermark proves nothing. Re-encoding or editing can strip a receipt, and watermark detection only covers the signals it supports.
  • The instrumental music generator available today produces output that is not cleared for commercial release, and the screen, the response and the embedded receipt all say so. Its commercially clear replacement is being built from scratch on rights-checked material and stays switched off until both of its stages are trained and a frozen evaluation passes without regression.
  • Voice enrollment currently offers five languages: US English, Mexican Spanish, French, German and Brazilian Portuguese.
  • Transcription accepts up to fifteen minutes or 18 MB per file; audio for analysis is capped at 50 MB; a speech render is capped at 2,500 characters and a scene at twelve lines.

Clone your voice without giving it away.

One audio workspace for voice, speech, transcripts, music, mixing and evidence, with no third-party AI service involved.

Prefer email? contact@autosecurelogin.com