原文地址:http://blog.csdn.net/wgw335363240/article/details/39889979
solr是基於 lucence開發的應用,如果query中帶有非法字符串,結果很可能是檢索出所有內容或者直接報錯,所以你對用戶的輸入必須要先做處理。輸入星號,能夠檢索出所有內容;輸入加號,則會報錯。
官方的處理辦法(java,因為solr是java開發的):
https://svn.apache.org/repos/asf/lucene/dev/trunk/solr/solrj/src/java/org/apache/solr/client/solrj/util/ClientUtils.java public static String escapeQueryChars(String s) { StringBuilder sb = new StringBuilder(); for (int i = 0; i < s.length(); i++) { char c = s.charAt(i); // These characters are part of the query syntax and must be escaped if (c == '\\' || c == '+' || c == '-' || c == '!' || c == '(' || c == ')' || c == ':' || c == '^' || c == '[' || c == ']' || c == '\"' || c == '{' || c == '}' || c == '~' || c == '*' || c == '?' || c == '|' || c == '&' || c == ';' || c == '/' || Character.isWhitespace(c)) { sb.append('\\'); } sb.append(c); } return sb.toString(); }
翻譯的php版本(利用preg_replace函數進行正則替換):
static public function escape($value) { //list taken from http://lucene.apache.org/java/docs/queryparsersyntax.html#Escaping%20Special%20Characters $pattern = '/(\+|-|&|\||!|\(|\)|\{|}|\[|]|\^|"|~|\*|\?|:|;|~|\/)/'; $replace = '\\\$1'; return preg_replace($pattern, $replace, $value); }
翻譯后的python版本:
import re def escape_solr(word): return re.sub('(\\\|\+|-|&|\|\||!|\(|\)|\{|}|\[|]|\^|"|~|\*|\?|:|;|/|\~)','\\\1', word )
