Why asking since the last run goes wrong
The obvious design is to remember when the job last ran and ask for everything changed after that moment. It fails in two ways. A record can become visible after it was created: HubSpot's search documentation notes that new or changed records may take a few moments to appear, so a record stamped before the last run can first show up after it, and the next run, asking only for later changes, never sees it. And if the job asks again from the same point after a crash, or from an overlap you added to catch late records, it sees records it has already handled and acts on them a second time. The example page runs five variants over the same synthetic records so the effect is visible.
Four pieces that fix it
A checkpoint is the saved position, kept somewhere that survives a crash. An overlap means asking from a little before the checkpoint, so that late-appearing records are caught. A key is whatever uniquely identifies a record, used to recognise ones already handled so that the overlap does not cause repeats. Complete paging means reading every page of results until there are none, because stopping after the first page loses everything beyond it; Zapier's developer documentation says it does not automatically fetch additional pages, so only the first page of a polling trigger's results is effectively checked, and it expects the newest items there. Together these make the job tolerant of slow, repeated and interrupted data.
- Choose an overlap longer than the longest delay you can observe.
- Choose a key the API guarantees is unique, such as its own id.
- Stop only when the API says there are no more pages, or the documented cap is reached and logged.
Order of operations decides what a crash does
There are three steps for each record: act on it, record its key as handled, and, once the whole run is recorded, move the checkpoint. The checkpoint always goes last, so a crash leaves it behind and the next run reads the same records again. The question is the order of the first two, and each order has a failure.
If you record the key first and then act, a crash between the two leaves a key marked handled whose action never ran. The next run skips it, and the record is lost without any error. If you act first and then record the key, a crash between the two means the next run acts again. For an action that adds something each time it runs, such as sending an email or appending a row, that is a duplicate. The table shows one record and a crash at the one risky point, with the next run starting from the unmoved checkpoint.
Acting and recording touch two different systems, so no ordering makes them one step. What works is to act first, in a way that leaves one effect however often it runs for the same record key, then record the key. Create-or-update on the key does that, as does a destination that accepts an idempotency key. The repeated call after a crash then changes nothing. The other route is to write the result and the handled key to one store in one transaction. Where the action can be neither, the choice between a possible repeat and a possible loss remains, and it has to be made on purpose and written down.
- Never advance the checkpoint before the work for that page is recorded.
- Act first in a way that is safe to repeat for the same key, then record the key.
- Keep handled keys for at least the overlap plus the longest gap between runs.
- Test by stopping the job before each step in turn, restarting it and counting the effects per key.
If the matrix is wider than the box, scroll horizontally to read every column. Keyboard: focus the matrix and use Left/Right.
One record K1. A crash at the risky point, then the next run starts from the unmoved checkpoint.
order crash point action effects for K1
key first, then act after recording, before acting any 0 (lost)
act first, then key after acting, before recording adds each time 2 (duplicate)
act first, then key after acting, before recording safe to repeat 1
Effects are counted at the destination, not calls made.Know the API's limits before you design around it
Read the list endpoint's documentation first. HubSpot's search allows at most 200 objects per page, caps a query at 10,000 total results and limits the request rate, which matters for a long backlog. Zapier's polling guidance expects newest-first order and a unique id for each item. If the API cannot filter by modification time, or offers no stable order, a complete and repeat-free job may not be possible, and you should say so rather than build one that appears to work. Where the API offers webhooks, consider whether they fit better, remembering that they bring their own repeat and loss behaviour.
Retries and repeat safety
RFC 9110 says a client should not automatically retry a request with a non-idempotent method, unless it knows that the request is effectively idempotent or can detect that the original was never applied. A job that retries a create call after a timeout may therefore create twice. The remedy is to make the action idempotent by key, for example create-or-update on the record's id, or to look the record up before retrying. This is why the action should be safe to repeat for the same key even when the API itself is well behaved, and why the handled-key record alone is not enough.
What this guide does not cover
This guide does not apply to the built-in polling triggers of Zapier or Make, which manage their own checkpoints. It does not promise that a live API behaves as its documentation says, or that a record appearing later than the overlap is caught. The paid outcome builds one scheduled job against one documented API, with a checkpoint, an overlap, handled keys and complete paging, and proves it against a synthetic stand-in for 120 records over three pages, a late arrival and a stop before each step in turn. It needs the action to be safe to repeat for the same record key. Your team deploys and schedules it.
Sources and limits
- Zapier: polling trigger de-duplication Checked 2026-10-11.
- Polling triggers need a unique id for every item, should return results newest first, and Zapier stores the ids it has seen and clears the list when the Zap is turned off.
- Zapier does not automatically fetch additional pages, so effectively only the first page of results is checked.
- HubSpot: CRM search API Checked 2026-10-11.
- A search returns at most 200 objects per page, is limited to 10,000 total results per query, and new or changed records may take a few moments to appear.
- IETF RFC 9110, section 9.2.2 Checked 2026-10-11.
- A client should not automatically retry a request with a non-idempotent method unless it knows the request is effectively idempotent or can detect that the original was not applied.