The scariest thing in automation isn’t when it doesn’t run.
It’s when it looks like it’s running, but nothing is actually happening.
I built a system to auto-post article announcements to X (@YKStudioLab). Windows Task Scheduler runs it 3 times a day, working through queued text files one at a time. The logs looked normal. The posted/ folder had 19 posts neatly archived as sent.
When I opened the account, it said 0 posts. Not a single one had gone out.
How it works
Let me lay out the setup first.
x_posts/
queue/ 0022_site_launch.txt ← one per run, lowest number first
posted/ 2026-08-30-14-00-01_0021_xxx.txt ← moved here after posting
_hold/ unverified posts with images, parked here
It uses Playwright with an existing Chrome profile (already logged in), opens x.com/compose/post, types the text, and presses the post button. If the post goes through, it moves the file from queue/ to posted/. That’s all. A very simple setup.
Too simple, as it turned out.
Cause 1: Treating the fact of a click as proof that it was sent
The original code looked like this:
await postBtn.click({ force: true });
await page.waitForTimeout(4000);
await archivePost(file); // ← archives unconditionally
The assumption was that if click() doesn’t throw, it succeeded. But all click() guarantees is that “a click event was dispatched at those coordinates.” It guarantees nothing about whether X accepted it and sent the post.
The queue gets consumed. The file moves to posted/. The next run moves on to the next article. So the failure never surfaced once, and 19 posts quietly evaporated.
The fix was just adding verifyPosted() so that the queue isn’t consumed unless the send can be confirmed. If it can’t be confirmed, the file stays in queue/, so the next scheduled run retries it automatically.
Cause 2: A “just in case” Escape was closing the compose modal
Before typing, I was unconditionally pressing Escape to close any popups that might be in the way.
await page.keyboard.press('Escape'); // close modals etc. (supposedly)
/compose/post is itself a modal. That Escape closed the compose dialog and dropped back to the home timeline behind it. The text field it then found to type into was the inline composer on the home timeline.
In other words, every single time, it wrote the post into the home composer and just left it there. It didn’t even become a draft.
I changed it to press Escape only when the input field can’t be found.
Cause 3: The button doesn’t respond even with force: true
X’s post button sometimes doesn’t respond even when you hit it with click({ force: true }). I gave up on reasoning through why and added a fallback.
await postBtn.click({ force: true }).catch(() => {});
// If still not sent after waiting 10 seconds, retry with X's post shortcut
if (!(await postedSoon(page, 10_000))) {
await page.keyboard.press('Control+Enter');
}
What actually got the post through was the Control+Enter.
What burned the most time was a bug in the fix
The first implementation of verifyPosted() checked the URL.
// This always says "success"
return !page.url().includes('/compose/post');
The idea was that on a successful post, the modal closes and the URL goes back to home. It sounds plausible at first glance.
But it’s a modal, so the URL goes back to home even when it fails and closes. Success and failure were completely identical in the value I was checking. The test right after I added this fix reported “posted successfully” and actually posted nothing. I had dug the same hole all over again, in the code meant to fix the cause.
Only one thing turned out to be reliable:
Has the post text disappeared from the input field?
Once a post is sent, the input field goes empty. On success there’s also a Your post was sent. toast, but it disappears quickly and is easy to miss, so I made the input field’s contents the primary check.
A DNS blip I hit along the way
On August 31, all 3 runs, at 9:00, 14:00, and 20:00, failed with net::ERR_NAME_NOT_RESOLVED.
I assumed it was an outage, so right afterward I ran a diagnostic tool with the same profile: x.com and twitter.com both returned HTTP 200, and the auth cookies were still valid. It wasn’t a lasting outage, just a momentary DNS blip.
Since it runs unattended, a single blip wastes that entire run. I added retries.
async function gotoWithRetry(page, url, tries = 4) {
for (let i = 0; i < tries; i++) {
try { return await page.goto(url, { waitUntil: 'domcontentloaded' }); }
catch (e) {
if (!/ERR_NAME_NOT_RESOLVED|ERR_NETWORK|ERR_CONNECTION/.test(String(e))) throw e;
if (i === tries - 1) throw e;
await page.waitForTimeout(5000 * (i + 1)); // 5s → 10s → 15s
}
}
}
It retries only network errors and rethrows everything else immediately. If you retry everything, real bugs turn into “it fails once in a while” and disappear from view, so I kept the condition narrow.
One more thing: the login check needs retries too. If a blip makes the login page fetch fail, it wrongly concludes “not logged in,” and in an unattended run it waits 5 minutes and then fails. The check runs before the main process, so if it’s fragile, all the hardening of the main process is wasted.
Now
I made a real post of 0022_site_launch and confirmed everything down to the OGP card rendering. 5 posts remain in the queue. The 2 with images are still parked in _hold/, and I haven’t verified the media-attachment path with a real post yet. I’ve decided not to write “it probably works” about paths I haven’t verified, so I’m marking this one as unverified.
The 19 files in posted/ never actually went out. If the content hasn’t gone stale, they’re worth reposting.
Lessons
Treat “it ran” and “it had an effect” as two separate things. Clicks, API calls, file moves: these are all records of execution, not records of results. The condition for consuming an automation queue should have been confirming the result, not execution.
And one more: never check a value that comes out the same on success and failure. The URL check fell right into this. Asking myself once before writing it, “what will this value be when it fails?”, would have prevented it.
In the same “broken without any error” category, the time billing wasn’t working for 11 days did far more damage. I run both development and operations, this automation included, with the work split among AI agents. I wrote about that setup here.