@smart-cloud/publisher-exporter
v1.1.83
Published
Headless Playwright static publisher for WordPress/Elementor sites with sitemap-only page discovery, strict asset capture, escaped URL rewrite, structured logs, and targeted retry modes.
Downloads
5,634
Readme
@smart-cloud/publisher-exporter
@smart-cloud/publisher-exporter is the standalone Node.js CLI package behind WP Suite Static Publisher crawl, deploy, invalidate, targeted content-sync, and queue-runner workflows.
It is designed to run outside WordPress. The WordPress plugin only manages runtime config, queue state, and logs.
Release notes
1.1.82
- Buffer cross-account S3 upload streams into 64 KiB chunks before the AWS SDK applies streaming checksums. This prevents real HTML and asset bodies from failing with
Only the last chunk is allowed to have a size less than 8192 byteswhile preserving bounded streaming instead of loading complete objects into Lambda memory. - Keep deploy-only recovery on the existing remote export manifest: after updating delegated workers, rerun only the deploy job to copy the already rendered and rewritten workspace objects without another crawl.
- Make relative URL rewriting atomic and repeatable. Overlapping absolute, root-relative, escaped, and protocol-relative rules no longer rewrite output produced by an earlier rule or add another
../prefix when a file passes through more than one rewrite phase. - Remove nested browser-runtime DOM before serializing rendered pages, not only top-level runtime roots. Authored placeholders and bootstrap scripts remain, so React, Mantine, Stripe, and similar client integrations initialize exactly once on the published page instead of duplicating pre-rendered controls.
Upgrade the orchestrator and delegated workers to 1.1.82. A WordPress plugin update is not required. A deploy-only retry is sufficient for the cross-account S3 upload failure; sites affected by malformed relative URLs or serialized runtime UI need a full publish so every affected HTML page is rewritten or rendered from its authored source.
1.1.81
- Let the base deployment target and each named deployment profile select an optional AWS shared-config profile for local S3 deployment and CloudFront invalidation. One-time credentials attached to a queued job still take precedence, and only the profile name is stored in Publisher configuration.
- Let
remote-workers.jsonindependently select a host AWS profile when the coordinator must reach a worker stack through a named identity. This worker-plane profile is separate from the active deployment target profile. - Allow an explicitly allowlisted delegated deploy target to assume a target-account role with an optional external ID. The worker reads the source workspace with its execution role and streams the object to S3 with target-role credentials, so neither identity needs access to both buckets.
- Validate profile names, target role ARNs, and external IDs before creating AWS clients, while keeping target role details out of generated site runtime configuration.
Upgrade the external coordinator, redeploy the Lambda worker stack with the cross-account target allowlist, and reinstall its generated remote-workers.json before selecting that target. Sites using only same-account ambient credentials retain their existing behavior.
1.1.80
- Remove client-created top-level provider, portal, and overlay roots before serializing rendered HTML, including roots mounted through a bundled direct
react-dom/clientimport that is not observable throughwindow.ReactDOM. - Ensure published pages run their parser-authored bootstrap and create fresh client providers instead of mistaking serialized, listener-free DOM for a live mount.
- Add a real headless Chromium regression test that reproduces the provider marker, captures the page, reloads the serialized HTML, and verifies successful provider reinitialization.
Upgrade both the external coordinator and delegated render-worker image to 1.1.80, then rerender affected HTML pages. This is an exporter-only fix; no WordPress plugin update is required.
1.1.79
- Let targeted content-sync trust pages already recorded in the verified baseline manifest even when the prior local render workspace has been cleaned up after a successful publish.
- Keep strict sitemap coverage checks for URLs that are absent from the verified baseline; newly published pages remain queued until a later normal publish establishes their deployable output.
This is a coordinator-only fix. Upgrade and restart the external queue runner; delegated Lambda render workers do not need an update for this release.
1.1.78
- Remove serialized runtime-only
data-*-bound,data-*-initialized,data-*-mounted, anddata-*-hydratedmarkers before saving rendered HTML. Event listeners and in-memory framework instances cannot survive HTML serialization, so retaining those markers can suppress clean initialization on the published origin. - Apply the guard through the shared serializer used by local rendering, delegated Lambda rendering, timeout fallback capture, and partial recovery capture.
- Add regression coverage for the generic runtime-marker sanitizer and shared local/remote capture path.
Upgrade both the external coordinator and every delegated render-worker image to 1.1.78, then run a full publish. Existing static HTML retains stale runtime markers until the affected pages are rendered and deployed again.
1.1.77
- Report delegated rewrite progress as monotonic job-wide file and completed-batch totals, even when concurrent Lambda batches finish out of order.
- Keep per-task DynamoDB rewrite snapshots available at debug level instead of allowing their independent counters to overwrite the visible job activity.
- Apply the same aggregate progress reporting to deployment-profile rewrites.
This is a coordinator-only observability update. Upgrade the external coordinator; no WordPress plugin or Lambda worker update is required.
1.1.76
- Restore parser-authored React mount points and shadow hosts to their pristine light-DOM state immediately before HTML serialization. Runtime detection uses
ReactDOM.createRoot()andattachShadow()browser primitives rather than plugin names or component-specific selectors. - Exclude script-created React providers, portals, scripts, stylesheets, styles, and iframes from static output while retaining their authored bootstrap code. Client runtimes therefore initialize from a clean state on the published origin instead of inheriting incomplete mount markers from the source render.
- Remove transient consent and reCAPTCHA UI before capture. This prevents a source-origin reCAPTCHA iframe and its hidden provider state from being frozen into a static page and producing timeouts after deployment.
- Apply the same serialization path to local rendering, delegated Lambda rendering, timeout fallback capture, and partial recovery capture.
Upgrade both the external coordinator and any delegated render-worker image to 1.1.76, then run a new full publish. Existing static HTML retains its captured runtime state until it is rendered and deployed again.
1.1.75
- Replace inherited remote-workspace objects when an incremental S3-native crawl produces a changed coordinator-owned output, including sitemap XML files.
- Keep Lambda outputs created by the current job authoritative, so local files cannot overwrite freshly rendered pages or remotely fetched assets.
- Add regression coverage for a changed sitemap inherited from a previous remote export manifest and for current-job object preservation.
Upgrade the external coordinator before the next incremental publish. Existing stale sitemap objects are not repaired automatically; run a new incremental publish after upgrading. Pinning the same exporter version in the Lambda worker image is supported for release parity, but this fix changes coordinator staging only.
1.1.74
- Add an opt-in, fail-closed WordPress page-cache purge preflight for crawl-like jobs. The coordinator authenticates with the existing private runtime token and accepts only an explicit complete-purge confirmation from the source site.
- Run the purge after invalidating the deploy plan and marking the crawl incomplete, but before deleting or writing export output. Authentication, transport, redirect, status, response-shape, or provider failures stop the crawl and preserve the previous files.
- Cover publish, crawl, retry-timeouts, single-URL, and content-sync jobs that enter rendering. Deploy-only, CDN-invalidation-only, and rewrite-resume operations do not purge.
- Keep the exporter cache-vendor-neutral. WordPress delegates to a versioned provider contract supplied by the host integration; the exporter contains no WP Super Cache, Cache Enabler, Nginx, or hosting-vendor implementation.
The preflight is disabled by default. Enable it only after installing and testing a source-site adapter. A complete purge makes the first request for every unique URL a cache miss; it guarantees freshness and repopulates the cache for later visitors or retries, but it is not a separate cache-warming phase.
1.1.73
- Retry page-render responses with HTTP 500–599 instead of silently skipping them. The coordinator applies a bounded 5/10/20/30-second backoff without increasing worker concurrency. Local rendering allows three attempts; delegated rendering retains
remoteWorkers.render.maxAttempts(default two, supported range one to five). - Retry only failed URLs in a remote batch, retain successful pages, and never register a 5xx error document as exported HTML. Local retries close their failed page/context before waiting and create a fresh context for the next attempt.
- Record exhausted HTTP retries as errors with URL, status and attempt count. An unresolved page 5xx leaves the crawl incomplete and blocks rewrite/deploy-plan creation and subsequent deployment, including full publishes.
- Preserve normal 4xx skip/tombstone handling and generated 404 capture. Asset/sitemap HTTP policies and
--retry-timeoutsselection are unchanged; this patch addresses page-render responses. - Add mixed-batch, retry-budget, error-body, cleanup, progress-counter and incomplete-export regression coverage.
Run npm run verify:page-http-retry for an isolated local HTTP-server/headless-Chromium integration test, or npm run build:premium && npm run verify:page-http-retry -- --built to test the distributable CLI. The check makes no AWS or WordPress production changes.
Upgrade the external coordinator after active jobs finish. Existing skipped pages are not repaired automatically: crawl the affected URLs again or run a new publish. This is a coordinator-only change; no Lambda browser configuration, handler or infrastructure change is required. The existing retry-timeouts mode does not select HTTP 5xx errors.
1.1.72
- Correct the coordinator's saved-page counter for HTML already uploaded to the S3 workspace by Lambda workers.
- Include validated retained local or remote outputs when incremental processing reuses a page, and when a completed checkpoint is resumed. Count each output path once per run, including repeated writes or retries.
- Keep failed requests, rejected URLs, missing outputs, and unrelated skips out of the saved-page count. Rendering, reuse decisions, asset processing, and deployment behavior are unchanged.
- Add behavioral regression coverage for local and S3-native saves, reuse, checkpoints, failures, and duplicate output accounting.
The counter correction takes effect in new coordinator jobs after upgrading the exporter. It does not modify an already running job or retrospectively rewrite its progress log.
1.1.71
- Restore
--no-zygotefor Lambda Chromium launches while keeping--single-processdisabled. Browser and renderer processes remain separate. - Retain one reusable browser per render batch and a fresh, non-persistent browser context for each page. Cookies and browser storage do not carry over between pages; the browser closes and its owned temporary files are cleaned at the batch boundary.
- Test the actual Playwright Headless Shell instead of substituting full Chromium. The lifecycle check runs three batches of 20 pages, verifies context isolation, samples browser-tree memory and temporary storage, and invokes the production temporary-file cleanup between batches.
- Validate the corrected launch configuration in an isolated ARM64 Lambda: 62 rendered pages across 11 invocations in the same warm execution environment. Local smoke tests remain additional checks, not a substitute for Lambda validation.
Upgrade the Lambda image's pinned exporter version to 1.1.71 and redeploy the worker stack. Updating only the WordPress coordinator's npm package does not update existing Lambda images. No architecture change is required.
Install
Global install:
npm install -g @smart-cloud/publisher-exporter
publisher-exporter install-browsersWithout a global install:
npx @smart-cloud/publisher-exporter install-browsersIf multiple users may run jobs on the same host, prefer a shared Playwright browsers path:
export PLAYWRIGHT_BROWSERS_PATH=/var/lib/playwright-browsers
publisher-exporter install-browsersIf PLAYWRIGHT_BROWSERS_PATH points to a shared system directory, create that directory first and make it writable by the same OS user that will run publisher-exporter install-browsers. A one-time elevated setup step to create or re-own the directory is fine. The later cron job does not need elevated privileges to use an already installed shared browser cache, but it does need read and execute access to that directory tree.
Commands
- Normal usage:
PUBLISHER_CONFIG=./publisher.config.json publisher-exporter crawl
PUBLISHER_CONFIG=./publisher.config.json publisher-exporter crawl --crawl-mode incremental
PUBLISHER_CONFIG=./publisher.config.json publisher-exporter crawl --resume-rewrite
PUBLISHER_CONFIG=./publisher.config.json publisher-exporter crawl --retry-timeouts
PUBLISHER_CONFIG=./publisher.config.json publisher-exporter deploy
PUBLISHER_CONFIG=./publisher.config.json publisher-exporter invalidate
publisher-exporter publish
publisher-exporter queue-runner --runtime-dir /srv/site/runtime --max-jobs 1Equivalent npx usage:
PUBLISHER_CONFIG=./publisher.config.json npx @smart-cloud/publisher-exporter crawl
PUBLISHER_CONFIG=./publisher.config.json npx @smart-cloud/publisher-exporter crawl --crawl-mode incremental
PUBLISHER_CONFIG=./publisher.config.json npx @smart-cloud/publisher-exporter crawl --resume-rewrite
PUBLISHER_CONFIG=./publisher.config.json npx @smart-cloud/publisher-exporter deploy
PUBLISHER_CONFIG=./publisher.config.json npx @smart-cloud/publisher-exporter invalidate
npx @smart-cloud/publisher-exporter queue-runner --runtime-dir /srv/site/runtime --max-jobs 1Crawl Modes And Repair Workflows
crawlwithout extra flags performs a normal full crawl, asset download, and final text rewrite.crawl --crawl-mode incrementalreuses the existing crawl manifest when possible and rewrites only the files affected by re-crawled pages, text assets, or changed asset-map entries. After a successful unlimited discovery pass it also removes exporter-generated HTML that is no longer present in the current manifest and records exact deletion tombstones for the deploy step. Post-crawl copy-map destinations remain protected.- An incremental crawl now fails closed when the remote WP Suite subscription configuration cannot be loaded. It never silently falls back to a full crawl that clears the existing output tree.
crawl --resume-rewriteskips discovery, rendering, and asset download, then reruns the final text rewrite over the existing output tree using the current rewrite rules.crawl --retry-timeoutsretries timed-out URLs from the latest archived full crawl or publish log snapshot.
Use --resume-rewrite when the already exported output is otherwise valid but the rewrite logic changed or a previous crawl stopped during the rewrite phase. This is the fastest repair path for stale output caused by rewrite-only bugs such as escaped JSON replacement issues, protocol-relative asset URLs like //host/path.css, or an interrupted final rewrite.
If you need to repair existing exported HTML after a rewrite bug, prefer --resume-rewrite over a new incremental crawl. Incremental crawl only rewrites targeted files; --resume-rewrite reprocesses every text file in the current output.
During --resume-rewrite, the terminal can be quieter than a full crawl. Progress still updates in runtime/current-progress.json and the live crawl event snapshot under the configured log directory.
Crawl integrity and deployment safety
Every crawl creates an incomplete-crawl marker before it changes the output.
The marker is removed only after final rewrite and deploy-plan generation complete.
deploy refuses to upload while that marker exists, so stopping a crawl cannot
fall back to scanning and uploading a partial or unre-written output tree.
If a crawl was stopped during final rewrite, run crawl --resume-rewrite; for
an interrupted crawl at an earlier stage, run a new crawl. Do not delete the
marker manually to bypass this protection.
Direct CLI With Runtime-Managed Config
When the WordPress plugin already wrote runtime/config.json into shared publisher storage, direct CLI invocation still needs both the config path and the runtime directory.
Example:
STATIC_PUBLISHER_RUNTIME_DIR=/mnt/site/runtime \
PUBLISHER_CONFIG=/mnt/site/runtime/config.json \
publisher-exporter crawl --resume-rewritePUBLISHER_CONFIG selects the main exporter config file. STATIC_PUBLISHER_RUNTIME_DIR anchors runtime artifacts such as current-progress.json, crawl-manifest.json, queue files, and storage-relative output/log paths.
Queue Runner
The queue runner reads the runtime JSON files generated by the WordPress plugin.
Example:
publisher-exporter queue-runner \
--runtime-dir /var/www/site/wp-content/uploads/smartcloud-static-publisher/runtime \
--max-jobs 1Direct CLI invocation from cron is the recommended setup.
When the runtime config contains enabled scheduler rules and the active queue-runner policy allows scheduler auto-enqueue, queue-runner evaluates those rules once at startup before draining queued jobs.
- Scheduler rules are read from
runtime/config.jsonand their last interval buckets are persisted inruntime/scheduler-state.json. - Scheduler only auto-enqueues jobs into the runtime queue. It does not replace cron, systemd timers, or Windows Task Scheduler.
- A 1-minute external runner tick is the recommended cadence.
- Supported scheduled commands are
publish,crawl,deploy,invalidate,retry-timeouts,url, and Professional/Agency-onlycontent-sync. - Content sync requires a successful normal publish baseline and processes only its configured post types, multisite scope, listings, archives, and sitemap surfaces.
- Crawl-producing jobs submit the privacy-minimal
privacy/cookie-observations.jsonartifact to the authenticated WordPress bridge when a runtime token is configured; installed consent providers keep candidates pending manual classification. - Targeted content sync can render or tombstone exact plugin-owned public resources declared through the versioned WordPress resource-change contract.
- The scheduler timezone value is currently informational for operations context; interval matching is based on elapsed minute buckets checked at each queue-runner start.
- If an equivalent queued or running job already exists for the same command, crawl mode, deployment profile, and URL, that rule is skipped for the current interval bucket.
Same-host Linux cron example:
SHELL=/bin/bash
HOME=/home/<runner-user>
PATH=/home/<runner-user>/.nvm/versions/node/v24.15.0/bin:/usr/bin:/bin
PLAYWRIGHT_BROWSERS_PATH=/var/lib/playwright-browsers
RUNTIME_PATH=/var/www/site/wp-content/uploads/smartcloud-static-publisher/runtime
LOG_PATH=/var/www/site/wp-content/uploads/smartcloud-static-publisher/logs
* * * * * /usr/bin/flock -n /tmp/static-publisher.cron.lock publisher-exporter queue-runner --runtime-dir "$RUNTIME_PATH" --max-jobs 1 >> "$LOG_PATH/queue-runner-cron.log" 2>&1
17 3 * * * publisher-exporter prune-logs --runtime-dir "$RUNTIME_PATH" --older-than-days 30 >> "$LOG_PATH/prune-logs-cron.log" 2>&1If WordPress and the runner are on different machines but share the same mounted publisher storage, keep outputDir and logDir storage-relative in WordPress admin and point cron at the crawler host's local mount path:
SHELL=/bin/bash
HOME=/home/<runner-user>
PATH=/home/<runner-user>/.nvm/versions/node/v24.15.0/bin:/usr/bin:/bin
PLAYWRIGHT_BROWSERS_PATH=/var/lib/playwright-browsers
RUNTIME_PATH=/mnt/site/runtime
LOG_PATH=/mnt/site/logs
* * * * * /usr/bin/flock -n /tmp/static-publisher.cron.lock publisher-exporter queue-runner --runtime-dir "$RUNTIME_PATH" --max-jobs 1 >> "$LOG_PATH/queue-runner-cron.log" 2>&1If you do not want a version-pinned NVM path in crontab, create a stable user launcher in ~/bin and put that directory first in PATH:
mkdir -p "$HOME/bin"
cat > "$HOME/bin/publisher-exporter" <<'EOF'
#!/usr/bin/env bash
set -euo pipefail
export HOME=/home/<runner-user>
export NVM_DIR="$HOME/.nvm"
. "$NVM_DIR/nvm.sh"
nvm use default >/dev/null
exec "$(npm prefix -g)/bin/publisher-exporter" "$@"
EOF
chmod +x "$HOME/bin/publisher-exporter"If cron already sets the correct HOME, remove the explicit export HOME=... line from the wrapper. A plain symlink to ~/.nvm/versions/node/vX.Y.Z/bin/publisher-exporter will break after upgrades. Prefer this small launcher, or enable an NVM-managed current symlink and link against that stable path.
After each finished, failed, or stopped job, queue-runner copies the exporter-generated working logs into "<logDir>/archive/<timestamp-command-jobId-status>/" and includes a job.json summary plus the latest current-progress.json snapshot when available. The live root log files remain the current working set and may be overwritten by the next job.
Archived files are gzip-compressed per file and listed in job.json. WordPress Audit Log rows can download those archived artifacts through the authenticated REST API.
retry-timeouts uses the newest archived full crawl or publish job log snapshot as its timeout source. It prefers the manifest-backed archived errors.* artifact when available and falls back to older uncompressed archive layouts.
Prune old archive directories with publisher-exporter prune-logs --runtime-dir /srv/site/runtime --older-than-days 30. Add that command to daily cron or another retention scheduler that matches your log policy.
AWS credential profiles
The base deployment target and each named deployment target may set an optional
awsProfile. Local S3 deploy and CloudFront invalidation resolve that profile
from the queue-runner user's standard AWS shared config. Store only the profile
name in publisher configuration; never store access keys there. Explicit
temporary credentials attached to a queued job take precedence over the named
profile.
This setting is deliberately host-local. A delegated Lambda deploy cannot read the WordPress VM's AWS profiles. Its CDK target uses either the deploy worker execution role or an allowlisted cross-account target role. CloudFront invalidation still runs on the coordinator and therefore uses the selected host profile when one is configured.
Optional Lambda workers
The exporter includes protocol-v1 building blocks for delegating bounded
render, asset-fetch, text-rewrite, and S3 deploy-copy tasks to container-image Lambda
functions. Large payloads do not travel in the Lambda response: workers write
HTML, assets, and detailed results below a configured S3 workspace prefix and return a
small resultKey pointer. Reusing the same job/task ID is idempotent.
The companion CDK application is maintained separately at
smartcloudsol/static-publisher-lambda-workers.
Its deploy script emits a remote-workers.json containing the
versioned Lambda alias ARNs, workspace bucket/prefix, region, and generated
assumable role. The file contains no static AWS access keys.
When the worker stack itself is reached through a host named profile, the file
may also contain awsProfile. That profile selects credentials for Lambda,
workspace S3, and progress DynamoDB access and is independent of the deployment
target profile. Without it, the exporter assumes the generated roleArn from
ambient credentials as before.
Install that output with owner-only permissions at
<site-runtime>/remote-workers.json, then validate the role, aliases, and
function runtimes:
publisher-exporter remote-worker-health \
--runtime-dir /srv/site/runtimeProtocol v1 accepts 1-20 same-origin URLs per render task. The default render batch size is 5, so a warm Chromium process can render several pages sequentially in one invocation while Lambda concurrency supplies parallelism between batches. Only failed or deferred URLs are retried. The Lambda response is not the rendered body; read the returned S3 result object using the same temporary role credentials.
Enable unified Lambda processing in the Static Publisher WordPress admin. The
single Processing concurrency value is dual-written to the legacy per-phase
fields for mixed-version compatibility. In unified mode, HTML and asset bodies
remain in the workspace through rewrite and direct S3-to-S3 deployment; the
coordinator transfers only task/result metadata. Older mixed phase settings
remain supported until an administrator changes the common controls. A current
remote-workers.json must contain render, asset, rewrite, and deploy-copy aliases.
When it also contains status.tableName, the exporter reads strongly consistent
DynamoDB heartbeats while a synchronous Lambda invocation is running and
forwards them to the existing runtime progress snapshot. Workers write those
rows directly with conditionally increasing sequence numbers.
Workers prefer the origin response Content-Type, fall back to the resource
extension only when the header is absent or invalid, and preserve that value
when writing transformed HTML, CSS, JavaScript, or other text back to S3. The
deploy worker includes Content-Type in its unchanged-object decision. After
upgrading from an earlier delegated rewrite implementation, republish affected
targets to replace already incorrect object metadata.
Each render invocation starts one normal multi-process Chromium instance and
creates a fresh non-persistent Playwright BrowserContext for every page. Closing
that context discards the page's cookies, cache, storage, routes, and child
pages without restarting Chromium; the browser is closed only after the batch.
Browser profiles and writable browser state are confined to the owned
/tmp/wpsuite-publisher-browser root. Workers reap orphaned Chromium processes
and clean only that root before launch and after the batch. If Chromium itself
disconnects during a page render, the worker starts a clean browser and retries
that URL once.
Structured Remote worker checkpoint. records include the request, job, task,
page, attempt, remaining Lambda time, process memory, /tmp and /dev/shm
space, and the visible Playwright profile count. Lambda Runtime.ExitError
events are emitted by the platform after the runtime has already stopped, so
the last checkpoint before the matching REPORT line identifies where the
runtime disappeared. This CloudWatch Logs Insights query keeps both records in
one timeline:
fields @timestamp, @requestId, @message
| filter @message like /Remote worker checkpoint|Remote worker runtime diagnostic|Runtime.ExitError/
| sort @timestamp ascThe saved runtime configuration remains compatible with older exporters:
{
"lambdaDelegationEnabled": true,
"processingConcurrency": 4,
"remoteWorkers": {
"render": {
"enabled": true,
"concurrency": 4,
"batchSize": 5,
"maxAttempts": 2
},
"rewrite": {
"enabled": true,
"concurrency": 4,
"batchSize": 25,
"maxAttempts": 2
},
"deploy": {
"enabled": true,
"concurrency": 4,
"batchSize": 50,
"maxAttempts": 2
}
}
}The normal queue runner already supplies the runtime directory, so cron needs no Lambda-specific environment variables. A one-off diagnostic crawl can still override the generated file explicitly:
STATIC_PUBLISHER_RUNTIME_DIR=/srv/site/runtime \
publisher-exporter crawl --remote-renderUse --local-render only as a legacy render diagnostic override; it disables
the unified remote data path for that crawl. The coordinator keeps sitemap
discovery, incremental and safety decisions, compact manifests, and deploy
planning local. Admitted pages are grouped into bounded Lambda batches. Their
HTML stays in S3, while each result returns discovered pages/assets and an
object reference. Asset workers fetch and inspect bounded URL batches. Final text
objects are rewritten in place by Lambda, then deploy workers check or copy
them directly from the workspace to the allowlisted target. Only explicit
deploy-plan tombstones may be deleted. Each completed
batch updates phase-scoped runtime progress and adds a structured batch summary
to deploy.log.jsonl; debug logging also records each returned object result.
There is no silent local fallback for any
enabled remote phase.
Troubleshooting
- If exported HTML still contains old rewritten URLs after a rewrite fix, run
crawl --resume-rewrite. This reprocesses every text file in the current output tree and does not depend on earlierrewriteComplete: trueentries incrawl-manifest.json. - If
--resume-rewriteappears to stop after the initial resume message, checkruntime/current-progress.jsonand<logDir>/current-crawl-event.jsonbefore assuming it is stuck. That phase can stay mostly quiet on stdout while it rewrites files. - If direct CLI execution seems to ignore the expected runtime state, verify that
PUBLISHER_CONFIGandSTATIC_PUBLISHER_RUNTIME_DIRboth point at the same runtime tree. A mismatched config path and runtime dir can make the run look idle or incomplete. - Prefer incremental crawl when the existing manifest is trusted and you only need a fast recrawl of changed pages or assets. Prefer a full crawl when discovery rules changed, the manifest is missing or suspect, or you want a clean end-to-end rebuild.
- JavaScript bundle discovery follows only explicit dynamic
import()calls and fixed Webpack chunk-filename mappings such asruntime.u = () => "custom-block-parser.js". It does not infer dependencies from arbitrary quoted paths or application-specific loader calls; add those manually as extra assets when needed. - In
sdk-upload-deletemode, remote-only S3 objects are protected unless the crawl deploy plan contains an explicit manifest-backed deletion tombstone. Successful unlimited incremental discovery reconciles exporter-generated local HTML with the current crawl manifest and creates those exact tombstones for orphaned HTML. Protected keys are listed indeploy-diff.jsonand the deploy log. This intentionally favors availability over automatically removing an unverified remote-only object. queue-runner-heartbeat.jsonreports both the exporter package version and the installed@smart-cloud/wpsuite-coreversion. Crawl logs also report the runtimevirtualAssetBaseUrl, remote config loader status, and resolved subscription type without logging the runtime token. Legacy runtime input usinguploadUrlis accepted during migration but normalized tovirtualAssetBaseUrl.- For generated 404 capture, set
generated404RequestPathto a path that really returns HTTP 404 from the origin. The exporter validates that response before reusing it as the generated 404 page.
WordPress Integration
The companion WordPress plugin can optionally store an External exporter dir setting. Point it at this package root when you want PHP-side diagnostics to verify the local CLI install.
Typical values:
/usr/local/lib/node_modules/@smart-cloud/publisher-exporter/opt/smartcloud/publisher-exporter/node_modules/@smart-cloud/publisher-exporter
Notes
- Playwright browser binaries are separate from the npm package and may need reinstalling after Playwright version upgrades.
- For internal/self-signed TLS origins, set
ignoreHttpsErrorsinpublisher.config.jsonor through the WordPress admin UI. - The package expects Node.js 20 or newer.
