3 ms·
I run a company that integrates data from hundreds of sources, including Shopify. There are only about 5 simple mistakes that most data extraction APIs mess up,
by slap_shot 7y ago
I run a company that integrates data from hundreds of sources, including Shopify. There are only about 5 simple mistakes that most data extraction APIs mess up, and this is one of them.
Very few APIs that implement pagination work optimally. If I query for all orders whose updated_at is greater than 2019-10-23 07:00:00, and paginate through the results, there's a good chance that any record updated before my pagination completes will be missed by the paginated queries. If I "checkpoint" the greatest updated_at retrieved in the most recent query, I will likely miss the records updated after my query started but before my query completed. Leaving me to use the start time that I began retrieving data as my new checkpoint.
With a cursor based pagination system, there is at least a chance the service that I'm calling to is dynamically adjusting their underlying query to account for this scenario.
Out of curiosity, what are the scenarios when you want to jump between pages (e.g. not iterate over the pages in a serial fashion)?
- QueensGambit 7y agoUsually, it is for syncing products or orders. If the API allows me to query by page number, I can execute all the pages in parallel and complete the job in seconds. If it is paginated by cursor, I have to do it one by one serially which might take minutes. Since I run this on serverless compute which has timeout (or browser timeout if user is waiting on it), splitting these sync jobs is needlessly complicated.
- QueensGambit 7y agoCan you please list the other 4 mistakes APIs make?
- slap_shot 7y agoOf course! I'm actively trying to standardize data extraction APIs, so I'm always happy to spread the good word. :) 1. The API should expose parameters for incrementally retrieving data by, at the very least, the date the record was created, the date the record was updated, and by a monotonic increasing value. The API should also expose parameters for sorting data by, again, at least the three fields mentioned earlier. These parameters should be consistent across _every_ endpoint. 2. Every record in the system should have a unique identifier. For instance, in Shopify, each record in order.tax_lines should have a unique identifier. Since they do not, you have to create a composite key of the order_id and its index, or take a hash from some set of values in the record. 3. Every entity in the system should be retrievable by the parameters exposed in rule #1. To get nested entities, many APIs require you fetch the "parent" entity, and then make a separate API call for each record to get child entity. For instance /customers might return 200 records, and then for each customer, you have to call /customer/{{customer_id}}/payment_sources. So you now have to make 201 requests to get all customers' payment_sources, when there should just be a /customer_payment_sources endpoint. 4. The API should implement a leaky bucket algorithm with known and documented rules, allowing for the caller to configure an exponential backoff that is optimized for that API. The API should have a maximum limit of time (e.g. 5 minutes) after which no requests are made that the API completely resets the bucket. Smaller issues, but still annoying: * represent _all_ timestamps in a common format (I think yyyy-MM-dd HH:mm:ss is best. I do not like unix timestamps since you cannot easily programmatically tell if it is a timestamp or just a large integer). * return results in JSON (still lots of XML being tossed around) * keep a change log that can be programmatically checked. We spend hours a week checking APIs to see if things have changed. One last thought: Webhooks are becoming standard for many APIs, which is great. However, Webhooks only expose data going forward. Many businesses build of years of valuable in the origin system that cannot be retreived without a proper extraction API. Those should be built first, then webhooks (imo). We've built some pretty awesome code to allow us to retrieve data from virtually anywhere, but if companies would just follow the rules above, it would be 10-100x easier for everybody involved.
- QueensGambit 7y agoThis is great! You should create an API usability guide! On webhooks, how does the caller (like Shopify) care if the transaction is completed without errors? Is there any message queue built specifically for webhooks?
- slap_shot 7y agoThe bullet points I listed were quick summaries of a larger internal wiki that we are building and working directly with data providers to implement. We'll definitely slim it down to a nice guide for developers to reference when beginning the journey of building one of these APIs. > On webhooks, how does the caller (like Shopify) care if the transaction is completed without errors? Most services have rules for the number of times they will attempt to execute any single request before stopping. Further, they typically have rules for how many times they will execute _any_ request before finally unregistering the webhook. Some will queue the events up and try for days, others will unregister the webhook after 5 failed attempts in a window.