Elasticsearch search handled some of these cases by introducing keyword field type, see 10061. With Solr, there are two alternatives:
- use analysed field and let Solr uninvert field when needed - such solution affects query/index performances because uninverting is not free and it also can consume significant portion of heap
- analyse field on client side and send values as multivalued field - in case of simple value standardisation like lowercasing, this can be acceptable solution (assuming you control the client), but in more complex cases this would mean having analysis configuration in two places.
Fortunately, Solr has many extension points that can be used to customise its behaviour.
One such point is update request processor chain. We are able to manipulate input doc before it reaches processor that actually does indexing. We can use update request processor to use value(s) from one field, parse them using some analysis chain and adds resulting values of some other field. That other field can be string field and we can enable docValues on that field. Unfortunately, Solr does not provide such processor OOTB and we have to write one.
solr-multivaluefield-processor
Base implementation can be found at solr-multivaluefield-processor. It does exactly as it was describe in previous paragraph: parse values from some field and store them as values of some target field. It does not modify source field. The code can be modified to replace values of existing field, or it can be combined with existing processors to remove existing field if needed. It takes three mandatory parameters: sourceField, destField and analysisType:
<updateRequestProcessorChain name="multivalue">
<processor class="com.od.bits.solr.update.processor.MultivalueFieldUpdateProcessorFactory">
<str name="sourceField">tags_txt</str>
<str name="destField">tags_ss</str>
<str name="analysisType">text_en</str>
</processor>
...
</updateRequestProcessorChain>
Test run processor
I am not a Docker expert, but I tend to use Docker to do quick tests. If you are not using docker, you can easily adjust instructions to test processor on an existing Solr setup.Clone and build processor
Code is placed on Github and Maven is used to build it:
git clone https://github.com/od-bits/solr-multivaluefield-processor.git
cd solr-multivaluefield-processor
mvn package
solr-multivalue-processor-1.0.jar in target
Start Solr on Docker
This example uses Solr 6.5. and core with basic_configs is used as a starting point
docker run --name multivalue -it -p8983:8983 solr:6.5
docker exec -it multivalue bin/solr create_core -c multivalue -d basic_configs
Configure custom update request processor chain
Copy created jar to some folder accessible by Solr - we will usecontrib folder
docker cp target/solr-multivaluefield-processor-*.jar multivalue:/opt/solr/contrib/
solrconfig.xml:
<!-- load extension -->
<lib dir="${solr.install.dir:../../../..}/contrib/" regex="solr-multivaluefield-processor-\d.*\.jar" />
<!-- define custom processor chain -->
<updateRequestProcessorChain name="multivalue">
<processor class="com.od.bits.solr.update.processor.MultivalueFieldUpdateProcessorFactory">
<str name="sourceField">tags_txt</str>
<str name="destField">tags_ss</str>
<str name="analysisType">text_en</str>
</processor>
<processor class="solr.LogUpdateProcessorFactory" />
<processor class="solr.DistributedUpdateProcessorFactory" />
<processor class="solr.RunUpdateProcessorFactory" />
</updateRequestProcessorChain>
<!-- use custom processor chain as default -->
<initParams path="/update/**">
<lst name="defaults">
<str name="update.chain">multivalue</str>
</lst>
</initParams>
docker cp src/test/resources/solrconfig.xml multivalue:/opt/solr/server/solr/multivalue/conf/
curl 'localhost:8983/solr/admin/cores?action=RELOAD&core=multivalue'
curl -XPOST 'localhost:8983/solr/multivalue/update?commit=true' -d '[{"id": 1, "tags_txt": "solr docvalues"}]'
tags_ss field.
{
"responseHeader":{
"status":0,
"QTime":0,
"params":{
"q":"*:*",
"indent":"true",
"wt":"json"}},
"response":{"numFound":1,"start":0,"docs":[
{
"id":"1",
"tags_txt":["solr docvalues"],
"tags_ss":["solr",
"docvalu"],
"_version_":1592828851126272000}]
}}
Conclusion
Uninverted field caches are frequent cause of Solr issues and results in running Solr on larger heap as a quick solution. It is important to be able to enable docValues on any field that is used for sorting or faceting. Solr made us choose between analysis and docValues, but now there is a simple way to choose both. It is not a perfect solution, but it should be enough until SOLR-8362 is resolved.Now it is up to you to set up your analysis or adjust processor to suite your needs!
Post a Comment