DBQ or Delete-By-Query

Why DBQ works well in dev env but often fails in production? You tested it with the production indexing load all worked ok, but when you do the same thing in production, indexing is slower, you get timeouts, replicas are falling behind and recovering and in some cases it there are OOMs. It seems like you hit some Solr bug but unfortunately, it is expected behaviour. The reason for this behaviour is Solr doing necessary steps to keep index in consistent state. Or more precisely, keeping eventual consistency of all replicas of the same shard. What are those steps?

What's behind DBQ

When Solr receives DBQ request, it first needs to forward that request to all shards' leaders. Once leader gets DBQ request, it needs to make sure deletes operation is performed on the same set of documents on both leader and all its replicas. In order to achieve that, it needs to block all updates make sure that all pending update requests are flushed to replicas. Once those prerequisites are ensured, it can safely forward DBQ request to its replicas. While the leader is waiting for replicas to finish, it can accept updates but cannot forward them to replicas. This can result in replicas falling behind and leader initiated replica recoveries.
Why it did not happen when you were testing you code in dev env? Well, in dev env you usually don't have replicas or you even run Solr in non cloud mode.

Alternative approach

DBQ does not play well with indexing and should not be run at the same time. In order to void having stability issues, DBQ should be run before indexing starts. In case you have to delete some documents at the same time, DBQ should be implemented on the client side:
  • Use cursor to return list of IDs, page by page
  • For each page, create batch of delete-by-id requests for each returned id
Luckily, delete-by-id does not suffer from the same issues as DBQ as it can use standard optimistic locking of a single document. Also, it is safe to delete docuemtns page by page as deleting does not affect cursor.

Takeaways

DBQ should be treated as a convenient development tool and not as something that should be used in standard data flows. Note that ES does not support DBQ at all and it seems that it would simplify things for Solr as well if they decided to drop it.
Another takeaway is that before going live, you should test used SolrCloud features with at least 2x2 collection since sharding and replication comes with an overhead.

Post a Comment