<table><tr><td style="">bruns added a comment.

</td><a style="text-decoration: none; padding: 4px 8px; margin: 0 8px 8px; float: right; color: #464C5C; font-weight: bold; border-radius: 3px; background-color: #F7F7F9; background-image: linear-gradient(to bottom,#fff,#f1f0f1); display: inline-block; border: 1px solid rgba(71,87,120,.2);" href="https://phabricator.kde.org/D11552">View Revision</a></tr></table><br /><div><div><blockquote style="border-left: 3px solid #8C98B8;

          color: #6B748C;

          font-style: italic;

          margin: 4px 0 12px 0;

          padding: 8px 12px;

          background-color: #F8F9FC;">

<div style="font-style: normal;

          padding-bottom: 4px;">In <a href="https://phabricator.kde.org/D11552#231330" style="background-color: #e7e7e7;

          border-color: #e7e7e7;

          border-radius: 3px;

          padding: 0 4px;

          font-weight: bold;

          color: black;text-decoration: none;">D11552#231330</a>, <a href="https://phabricator.kde.org/p/hein/" style="

              border-color: #f1f7ff;

              color: #19558d;

              background-color: #f1f7ff;

                border: 1px solid transparent;

                border-radius: 3px;

                font-weight: bold;

                padding: 0 4px;">@hein</a> wrote:</div>

<div style="margin: 0;

          padding: 0;

          border: 0;

          color: rgb(107, 116, 140);"><p>For the record though - a better way to do this is to use QTextBoundaryFinder which will operate e.g. on grapheme cluster boundaries. This still isn't super great for Chinese though. If you want to really-properly do it you'll end up depending on ICU and using its BreakIterator combined with dict-based support for Chinese, which isn't terribly fast however.</p></div>

</blockquote>

<p>There are a few implications here:</p>

<ul class="remarkup-list">

<li class="remarkup-list-item">splitting to much generates to unspecific terms, especially in case of full text indexing (Think of splitting a western language at character level, most texts likely contain almost the full alphabet. Same likely applies to Katakana with its about ~100 graphemes)</li>

<li class="remarkup-list-item">term generation at query and index time have to agree about what a term is, otherwise a search will likely return nothing. Changing the splitting at a later time will require reindexing all affected files</li>

<li class="remarkup-list-item">better splitting will cost some more time at index generation, but likely makes searching faster (additional time for term generation will be neglegible, but the search terms are less complex - e.g. "abc" instead of "a" AND "b" AND "c").</li>

</ul></div></div><br /><div><strong>REPOSITORY</strong><div><div>R293 Baloo</div></div></div><br /><div><strong>REVISION DETAIL</strong><div><a href="https://phabricator.kde.org/D11552">https://phabricator.kde.org/D11552</a></div></div><br /><div><strong>To: </strong>michaelh, hein<br /><strong>Cc: </strong>bruns, lbeltrame, Frameworks, alexeymin, cfeck, ashaposhnikov, michaelh, astippich, spoorun, nicolasfella, ngraham<br /></div>