10 Lempel-Ziv Compression
\(\newcommand{\footnotename}{footnote}\)
\(\def \LWRfootnote {1}\)
\(\newcommand {\footnote }[2][\LWRfootnote ]{{}^{\mathrm {#1}}}\)
\(\newcommand {\footnotemark }[1][\LWRfootnote ]{{}^{\mathrm {#1}}}\)
\(\let \LWRorighspace \hspace \)
\(\renewcommand {\hspace }{\ifstar \LWRorighspace \LWRorighspace }\)
\(\newcommand {\TextOrMath }[2]{#2}\)
\(\newcommand {\mathnormal }[1]{{#1}}\)
\(\newcommand \ensuremath [1]{#1}\)
\(\newcommand {\LWRframebox }[2][]{\fbox {#2}} \newcommand {\framebox }[1][]{\LWRframebox } \)
\(\newcommand {\setlength }[2]{}\)
\(\newcommand {\addtolength }[2]{}\)
\(\newcommand {\setcounter }[2]{}\)
\(\newcommand {\addtocounter }[2]{}\)
\(\newcommand {\arabic }[1]{}\)
\(\newcommand {\number }[1]{}\)
\(\newcommand {\noalign }[1]{\text {#1}\notag \\}\)
\(\newcommand {\cline }[1]{}\)
\(\newcommand {\directlua }[1]{\text {(directlua)}}\)
\(\newcommand {\luatexdirectlua }[1]{\text {(directlua)}}\)
\(\newcommand {\protect }{}\)
\(\def \LWRabsorbnumber #1 {}\)
\(\def \LWRabsorbquotenumber "#1 {}\)
\(\newcommand {\LWRabsorboption }[1][]{}\)
\(\newcommand {\LWRabsorbtwooptions }[1][]{\LWRabsorboption }\)
\(\def \mathchar {\ifnextchar "\LWRabsorbquotenumber \LWRabsorbnumber }\)
\(\def \mathcode #1={\mathchar }\)
\(\let \delcode \mathcode \)
\(\let \delimiter \mathchar \)
\(\def \oe {\unicode {x0153}}\)
\(\def \OE {\unicode {x0152}}\)
\(\def \ae {\unicode {x00E6}}\)
\(\def \AE {\unicode {x00C6}}\)
\(\def \aa {\unicode {x00E5}}\)
\(\def \AA {\unicode {x00C5}}\)
\(\def \o {\unicode {x00F8}}\)
\(\def \O {\unicode {x00D8}}\)
\(\def \l {\unicode {x0142}}\)
\(\def \L {\unicode {x0141}}\)
\(\def \ss {\unicode {x00DF}}\)
\(\def \SS {\unicode {x1E9E}}\)
\(\def \dag {\unicode {x2020}}\)
\(\def \ddag {\unicode {x2021}}\)
\(\def \P {\unicode {x00B6}}\)
\(\def \copyright {\unicode {x00A9}}\)
\(\def \pounds {\unicode {x00A3}}\)
\(\let \LWRref \ref \)
\(\renewcommand {\ref }{\ifstar \LWRref \LWRref }\)
\( \newcommand {\multicolumn }[3]{#3}\)
\(\require {textcomp}\)
\(\newcommand {\intertext }[1]{\text {#1}\notag \\}\)
\(\let \Hat \hat \)
\(\let \Check \check \)
\(\let \Tilde \tilde \)
\(\let \Acute \acute \)
\(\let \Grave \grave \)
\(\let \Dot \dot \)
\(\let \Ddot \ddot \)
\(\let \Breve \breve \)
\(\let \Bar \bar \)
\(\let \Vec \vec \)
\(\renewcommand {\vec }{\boldsymbol }\)
\(\newcommand {\Edge }{\ensuremath {\,\textemdash \,}}\)
\(\newcommand \Const [1]{\text {\textsf {#1}}}\)
\(\DeclareMathOperator {\lerp }{lerp}\)
\(\DeclareMathOperator {\bitlen }{bitlen}\)
\(\DeclareMathOperator {\sign }{sign}\)
\(\newcommand {\I }{\mathrm {i}}\)
\(\newcommand \AND {\mathbin {\&}}\)
\(\newcommand \OR {\mathbin {|}}\)
\(\newcommand \XOR {\mathbin {{}^{\wedge }}}\)
\(\newcommand \shl {\ll }\)
\(\newcommand \shr {\ggg }\)
\(\newcommand \asr {\gg }\)
\(\newcommand \NOT {\ensuremath {\mathord {\sim }}}\)
\(\newcommand {\isep }{\mathrel {{.}\,{.}}}\)
\(\newcommand {\Id }[1]{\mathit {#1}}\)
\(\newcommand {\const }[1]{\mathsf {#1}}\)
\(\newcommand {\algorithmname }[1]{\text {\textsc {#1}}}\)
\(\newcommand {\bits }[1]{\text {#1}}\)
\(\newcommand {\hexa }[1]{\mathtt {0x#1}}\)
\(\newcommand {\num }[1]{#1}\)
\(\newcommand {\qed }{\quad \square }\)
\(\newcommand {\idiv }[2]{\lfloor #1/#2\rfloor }\)
\(\newcommand \attribdot {\ensuremath {\mkern 1.5mu.\mkern 1.5mu}}\)
\(\newcommand \attribxr [2]{#1\attribdot \text {#2}}\)
\(\newcommand \attribir [2]{\Id {#1}\attribdot \text {#2}}\)
\(\newcommand \attribii [2]{\Id {#1}\attribdot \Id {#2}}\)
\(\newcommand \textsc [1]{#1}\)
\(\require {colortbl}\)
\(\let \LWRorigcolumncolor \columncolor \)
\(\renewcommand {\columncolor }[2][named]{\LWRorigcolumncolor [#1]{#2}\LWRabsorbtwooptions }\)
\(\let \LWRorigrowcolor \rowcolor \)
\(\renewcommand {\rowcolor }[2][named]{\LWRorigrowcolor [#1]{#2}\LWRabsorbtwooptions }\)
\(\let \LWRorigcellcolor \cellcolor \)
\(\renewcommand {\cellcolor }[2][named]{\LWRorigcellcolor [#1]{#2}\LWRabsorbtwooptions }\)
\(\newcommand {\tcbset }[1]{}\)
\(\newcommand {\tcbsetforeverylayer }[1]{}\)
\(\newcommand {\tcbox }[2][]{\boxed {\text {#2}}}\)
\(\newcommand {\tcboxfit }[2][]{\boxed {#2}}\)
\(\newcommand {\tcblower }{}\)
\(\newcommand {\tcbline }{}\)
\(\newcommand {\tcbtitle }{}\)
\(\newcommand {\tcbsubtitle [2][]{\mathrm {#2}}}\)
\(\newcommand {\tcboxmath }[2][]{\boxed {#2}}\)
\(\newcommand {\tcbhighmath }[2][]{\boxed {#2}}\)
\(\newcommand {\toprule }[1][]{\hline }\)
\(\let \midrule \toprule \)
\(\let \bottomrule \toprule \)
\(\def \LWRbooktabscmidruleparen (#1)#2{}\)
\(\newcommand {\LWRbooktabscmidrulenoparen }[1]{}\)
\(\newcommand {\cmidrule }[1][]{\ifnextchar (\LWRbooktabscmidruleparen \LWRbooktabscmidrulenoparen }\)
\(\newcommand {\morecmidrules }{}\)
\(\newcommand {\specialrule }[3]{\hline }\)
\(\newcommand {\addlinespace }[1][]{}\)
\(\newcommand {\LWRsubmultirow }[2][]{#2}\)
\(\newcommand {\LWRmultirow }[2][]{\LWRsubmultirow }\)
\(\newcommand {\multirow }[2][]{\LWRmultirow }\)
\(\newcommand {\mrowcell }{}\)
\(\newcommand {\mcolrowcell }{}\)
\(\newcommand {\STneed }[1]{}\)
\(\newcommand {\LWRldelimtwo }[1][]{\text {#1}~\LWRbigdelim }\)
\(\newcommand {\LWRldelimone }[2][]{\LWRldelimtwo }\)
\(\def \ldelim #1#2{\def \LWRbigdelim {#1}\LWRldelimone }\)
\(\newcommand {\LWRrdelimtwo }[1][]{\LWRbigdelim ~\text {#1}}\)
\(\newcommand {\LWRrdelimone }[2][]{\LWRrdelimtwo }\)
\(\def \rdelim #1#2{\def \LWRbigdelim {#1}\LWRrdelimone }\)
\(\let \symnormal \mathit \)
\(\let \symliteral \mathrm \)
\(\let \symbb \mathbb \)
\(\let \symbbit \mathbb \)
\(\let \symcal \mathcal \)
\(\let \symscr \mathscr \)
\(\let \symfrak \mathfrak \)
\(\let \symsfup \mathsf \)
\(\let \symsfit \mathit \)
\(\let \symbfsf \mathbf \)
\(\let \symbfup \mathbf \)
\(\newcommand {\symbfit }[1]{\boldsymbol {#1}}\)
\(\let \symbfcal \mathcal \)
\(\let \symbfscr \mathscr \)
\(\let \symbffrak \mathfrak \)
\(\let \symbfsfup \mathbf \)
\(\newcommand {\symbfsfit }[1]{\boldsymbol {#1}}\)
\(\let \symup \mathrm \)
\(\let \symbf \mathbf \)
\(\let \symit \mathit \)
\(\let \symsf \symsfit \)
\(\let \symtt \mathtt \)
\(\let \symbffrac \mathbffrac \)
\(\newcommand {\mathfence }[1]{\mathord {#1}}\)
\(\newcommand {\mathover }[1]{#1}\)
\(\newcommand {\mathunder }[1]{#1}\)
\(\newcommand {\mathaccent }[1]{#1}\)
\(\newcommand {\mathbotaccent }[1]{#1}\)
\(\newcommand {\mathalpha }[1]{\mathord {#1}}\)
\(\def\Alpha{\unicode{x1D6E2}}\)
\(\def\Beta{\unicode{x1D6E3}}\)
\(\def\Gamma{\unicode{x1D6E4}}\)
\(\def\Digamma{\mathit{\unicode{x03DC}}}\)
\(\def\Delta{\unicode{x1D6E5}}\)
\(\def\Epsilon{\unicode{x1D6E6}}\)
\(\def\Zeta{\unicode{x1D6E7}}\)
\(\def\Eta{\unicode{x1D6E8}}\)
\(\def\Theta{\unicode{x1D6E9}}\)
\(\def\Vartheta{\unicode{x1D6F3}}\)
\(\def\Iota{\unicode{x1D6EA}}\)
\(\def\Kappa{\unicode{x1D6EB}}\)
\(\def\Lambda{\unicode{x1D6EC}}\)
\(\def\Mu{\unicode{x1D6ED}}\)
\(\def\Nu{\unicode{x1D6EE}}\)
\(\def\Xi{\unicode{x1D6EF}}\)
\(\def\Omicron{\unicode{x1D6F0}}\)
\(\def\Pi{\unicode{x1D6F1}}\)
\(\def\Rho{\unicode{x1D6F2}}\)
\(\def\Sigma{\unicode{x1D6F4}}\)
\(\def\Tau{\unicode{x1D6F5}}\)
\(\def\Upsilon{\unicode{x1D6F6}}\)
\(\def\Phi{\unicode{x1D6F7}}\)
\(\def\Chi{\unicode{x1D6F8}}\)
\(\def\Psi{\unicode{x1D6F9}}\)
\(\def\Omega{\unicode{x1D6FA}}\)
\(\def\alpha{\unicode{x1D6FC}}\)
\(\def\beta{\unicode{x1D6FD}}\)
\(\def\varbeta{\unicode{x03D0}}\)
\(\def\gamma{\unicode{x1D6FE}}\)
\(\def\digamma{\mathit{\unicode{x03DD}}}\)
\(\def\delta{\unicode{x1D6FF}}\)
\(\def\epsilon{\unicode{x1D716}}\)
\(\def\varepsilon{\unicode{x1D700}}\)
\(\def\zeta{\unicode{x1D701}}\)
\(\def\eta{\unicode{x1D702}}\)
\(\def\theta{\unicode{x1D703}}\)
\(\def\vartheta{\unicode{x1D717}}\)
\(\def\iota{\unicode{x1D704}}\)
\(\def\kappa{\unicode{x1D705}}\)
\(\def\varkappa{\unicode{x1D718}}\)
\(\def\lambda{\unicode{x1D706}}\)
\(\def\mu{\unicode{x1D707}}\)
\(\def\nu{\unicode{x1D708}}\)
\(\def\xi{\unicode{x1D709}}\)
\(\def\omicron{\unicode{x1D70A}}\)
\(\def\pi{\unicode{x1D70B}}\)
\(\def\varpi{\unicode{x1D71B}}\)
\(\def\rho{\unicode{x1D70C}}\)
\(\def\varrho{\unicode{x1D71A}}\)
\(\def\sigma{\unicode{x1D70E}}\)
\(\def\varsigma{\unicode{x1D70D}}\)
\(\def\tau{\unicode{x1D70F}}\)
\(\def\upsilon{\unicode{x1D710}}\)
\(\def\phi{\unicode{x1D719}}\)
\(\def\varphi{\unicode{x1D711}}\)
\(\def\chi{\unicode{x1D712}}\)
\(\def\psi{\unicode{x1D713}}\)
\(\def\omega{\unicode{x1D714}}\)
\(\def\upAlpha{\unicode{x0391}}\)
\(\def\upBeta{\unicode{x0392}}\)
\(\def\upGamma{\unicode{x0393}}\)
\(\def\upDigamma{\unicode{x03DC}}\)
\(\def\upDelta{\unicode{x0394}}\)
\(\def\upEpsilon{\unicode{x0395}}\)
\(\def\upZeta{\unicode{x0396}}\)
\(\def\upEta{\unicode{x0397}}\)
\(\def\upTheta{\unicode{x0398}}\)
\(\def\upVartheta{\unicode{x03F4}}\)
\(\def\upIota{\unicode{x0399}}\)
\(\def\upKappa{\unicode{x039A}}\)
\(\def\upLambda{\unicode{x039B}}\)
\(\def\upMu{\unicode{x039C}}\)
\(\def\upNu{\unicode{x039D}}\)
\(\def\upXi{\unicode{x039E}}\)
\(\def\upOmicron{\unicode{x039F}}\)
\(\def\upPi{\unicode{x03A0}}\)
\(\def\upVarpi{\unicode{x03D6}}\)
\(\def\upRho{\unicode{x03A1}}\)
\(\def\upSigma{\unicode{x03A3}}\)
\(\def\upTau{\unicode{x03A4}}\)
\(\def\upUpsilon{\unicode{x03A5}}\)
\(\def\upPhi{\unicode{x03A6}}\)
\(\def\upChi{\unicode{x03A7}}\)
\(\def\upPsi{\unicode{x03A8}}\)
\(\def\upOmega{\unicode{x03A9}}\)
\(\def\itAlpha{\unicode{x1D6E2}}\)
\(\def\itBeta{\unicode{x1D6E3}}\)
\(\def\itGamma{\unicode{x1D6E4}}\)
\(\def\itDigamma{\mathit{\unicode{x03DC}}}\)
\(\def\itDelta{\unicode{x1D6E5}}\)
\(\def\itEpsilon{\unicode{x1D6E6}}\)
\(\def\itZeta{\unicode{x1D6E7}}\)
\(\def\itEta{\unicode{x1D6E8}}\)
\(\def\itTheta{\unicode{x1D6E9}}\)
\(\def\itVartheta{\unicode{x1D6F3}}\)
\(\def\itIota{\unicode{x1D6EA}}\)
\(\def\itKappa{\unicode{x1D6EB}}\)
\(\def\itLambda{\unicode{x1D6EC}}\)
\(\def\itMu{\unicode{x1D6ED}}\)
\(\def\itNu{\unicode{x1D6EE}}\)
\(\def\itXi{\unicode{x1D6EF}}\)
\(\def\itOmicron{\unicode{x1D6F0}}\)
\(\def\itPi{\unicode{x1D6F1}}\)
\(\def\itRho{\unicode{x1D6F2}}\)
\(\def\itSigma{\unicode{x1D6F4}}\)
\(\def\itTau{\unicode{x1D6F5}}\)
\(\def\itUpsilon{\unicode{x1D6F6}}\)
\(\def\itPhi{\unicode{x1D6F7}}\)
\(\def\itChi{\unicode{x1D6F8}}\)
\(\def\itPsi{\unicode{x1D6F9}}\)
\(\def\itOmega{\unicode{x1D6FA}}\)
\(\def\upalpha{\unicode{x03B1}}\)
\(\def\upbeta{\unicode{x03B2}}\)
\(\def\upvarbeta{\unicode{x03D0}}\)
\(\def\upgamma{\unicode{x03B3}}\)
\(\def\updigamma{\unicode{x03DD}}\)
\(\def\updelta{\unicode{x03B4}}\)
\(\def\upepsilon{\unicode{x03F5}}\)
\(\def\upvarepsilon{\unicode{x03B5}}\)
\(\def\upzeta{\unicode{x03B6}}\)
\(\def\upeta{\unicode{x03B7}}\)
\(\def\uptheta{\unicode{x03B8}}\)
\(\def\upvartheta{\unicode{x03D1}}\)
\(\def\upiota{\unicode{x03B9}}\)
\(\def\upkappa{\unicode{x03BA}}\)
\(\def\upvarkappa{\unicode{x03F0}}\)
\(\def\uplambda{\unicode{x03BB}}\)
\(\def\upmu{\unicode{x03BC}}\)
\(\def\upnu{\unicode{x03BD}}\)
\(\def\upxi{\unicode{x03BE}}\)
\(\def\upomicron{\unicode{x03BF}}\)
\(\def\uppi{\unicode{x03C0}}\)
\(\def\upvarpi{\unicode{x03D6}}\)
\(\def\uprho{\unicode{x03C1}}\)
\(\def\upvarrho{\unicode{x03F1}}\)
\(\def\upsigma{\unicode{x03C3}}\)
\(\def\upvarsigma{\unicode{x03C2}}\)
\(\def\uptau{\unicode{x03C4}}\)
\(\def\upupsilon{\unicode{x03C5}}\)
\(\def\upphi{\unicode{x03D5}}\)
\(\def\upvarphi{\unicode{x03C6}}\)
\(\def\upchi{\unicode{x03C7}}\)
\(\def\uppsi{\unicode{x03C8}}\)
\(\def\upomega{\unicode{x03C9}}\)
\(\def\italpha{\unicode{x1D6FC}}\)
\(\def\itbeta{\unicode{x1D6FD}}\)
\(\def\itvarbeta{\unicode{x03D0}}\)
\(\def\itgamma{\unicode{x1D6FE}}\)
\(\def\itdigamma{\mathit{\unicode{x03DD}}}\)
\(\def\itdelta{\unicode{x1D6FF}}\)
\(\def\itepsilon{\unicode{x1D716}}\)
\(\def\itvarepsilon{\unicode{x1D700}}\)
\(\def\itzeta{\unicode{x1D701}}\)
\(\def\iteta{\unicode{x1D702}}\)
\(\def\ittheta{\unicode{x1D703}}\)
\(\def\itvartheta{\unicode{x1D717}}\)
\(\def\itiota{\unicode{x1D704}}\)
\(\def\itkappa{\unicode{x1D705}}\)
\(\def\itvarkappa{\unicode{x1D718}}\)
\(\def\itlambda{\unicode{x1D706}}\)
\(\def\itmu{\unicode{x1D707}}\)
\(\def\itnu{\unicode{x1D708}}\)
\(\def\itxi{\unicode{x1D709}}\)
\(\def\itomicron{\unicode{x1D70A}}\)
\(\def\itpi{\unicode{x1D70B}}\)
\(\def\itvarpi{\unicode{x1D71B}}\)
\(\def\itrho{\unicode{x1D70C}}\)
\(\def\itvarrho{\unicode{x1D71A}}\)
\(\def\itsigma{\unicode{x1D70E}}\)
\(\def\itvarsigma{\unicode{x1D70D}}\)
\(\def\ittau{\unicode{x1D70F}}\)
\(\def\itupsilon{\unicode{x1D710}}\)
\(\def\itphi{\unicode{x1D719}}\)
\(\def\itvarphi{\unicode{x1D711}}\)
\(\def\itchi{\unicode{x1D712}}\)
\(\def\itpsi{\unicode{x1D713}}\)
\(\def\itomega{\unicode{x1D714}}\)
\(\let \lparen (\)
\(\let \rparen )\)
\(\newcommand {\cuberoot }[1]{\,{}^3\!\!\sqrt {#1}}\,\)
\(\newcommand {\fourthroot }[1]{\,{}^4\!\!\sqrt {#1}}\,\)
\(\newcommand {\longdivision }[1]{\mathord {\unicode {x027CC}#1}}\)
\(\newcommand {\mathcomma }{,}\)
\(\newcommand {\mathcolon }{:}\)
\(\newcommand {\mathsemicolon }{;}\)
\(\newcommand {\overbracket }[1]{\mathinner {\overline {\ulcorner {#1}\urcorner }}}\)
\(\newcommand {\underbracket }[1]{\mathinner {\underline {\llcorner {#1}\lrcorner }}}\)
\(\newcommand {\overbar }[1]{\mathord {#1\unicode {x00305}}}\)
\(\newcommand {\ovhook }[1]{\mathord {#1\unicode {x00309}}}\)
\(\newcommand {\ocirc }[1]{\mathord {#1\unicode {x0030A}}}\)
\(\newcommand {\candra }[1]{\mathord {#1\unicode {x00310}}}\)
\(\newcommand {\oturnedcomma }[1]{\mathord {#1\unicode {x00312}}}\)
\(\newcommand {\ocommatopright }[1]{\mathord {#1\unicode {x00315}}}\)
\(\newcommand {\droang }[1]{\mathord {#1\unicode {x0031A}}}\)
\(\newcommand {\leftharpoonaccent }[1]{\mathord {#1\unicode {x020D0}}}\)
\(\newcommand {\rightharpoonaccent }[1]{\mathord {#1\unicode {x020D1}}}\)
\(\newcommand {\vertoverlay }[1]{\mathord {#1\unicode {x020D2}}}\)
\(\newcommand {\leftarrowaccent }[1]{\mathord {#1\unicode {x020D0}}}\)
\(\newcommand {\annuity }[1]{\mathord {#1\unicode {x020E7}}}\)
\(\newcommand {\widebridgeabove }[1]{\mathord {#1\unicode {x020E9}}}\)
\(\newcommand {\asteraccent }[1]{\mathord {#1\unicode {x020F0}}}\)
\(\newcommand {\threeunderdot }[1]{\mathord {#1\unicode {x020E8}}}\)
\(\newcommand {\Bbbsum }{\mathop {\unicode {x2140}}\limits }\)
\(\newcommand {\oiint }{\mathop {\unicode {x222F}}\limits }\)
\(\newcommand {\oiiint }{\mathop {\unicode {x2230}}\limits }\)
\(\newcommand {\intclockwise }{\mathop {\unicode {x2231}}\limits }\)
\(\newcommand {\ointclockwise }{\mathop {\unicode {x2232}}\limits }\)
\(\newcommand {\ointctrclockwise }{\mathop {\unicode {x2233}}\limits }\)
\(\newcommand {\varointclockwise }{\mathop {\unicode {x2232}}\limits }\)
\(\newcommand {\leftouterjoin }{\mathop {\unicode {x27D5}}\limits }\)
\(\newcommand {\rightouterjoin }{\mathop {\unicode {x27D6}}\limits }\)
\(\newcommand {\fullouterjoin }{\mathop {\unicode {x27D7}}\limits }\)
\(\newcommand {\bigbot }{\mathop {\unicode {x27D8}}\limits }\)
\(\newcommand {\bigtop }{\mathop {\unicode {x27D9}}\limits }\)
\(\newcommand {\xsol }{\mathop {\unicode {x29F8}}\limits }\)
\(\newcommand {\xbsol }{\mathop {\unicode {x29F9}}\limits }\)
\(\newcommand {\bigcupdot }{\mathop {\unicode {x2A03}}\limits }\)
\(\newcommand {\bigsqcap }{\mathop {\unicode {x2A05}}\limits }\)
\(\newcommand {\conjquant }{\mathop {\unicode {x2A07}}\limits }\)
\(\newcommand {\disjquant }{\mathop {\unicode {x2A08}}\limits }\)
\(\newcommand {\bigtimes }{\mathop {\unicode {x2A09}}\limits }\)
\(\newcommand {\modtwosum }{\mathop {\unicode {x2A0A}}\limits }\)
\(\newcommand {\sumint }{\mathop {\unicode {x2A0B}}\limits }\)
\(\newcommand {\intbar }{\mathop {\unicode {x2A0D}}\limits }\)
\(\newcommand {\intBar }{\mathop {\unicode {x2A0E}}\limits }\)
\(\newcommand {\fint }{\mathop {\unicode {x2A0F}}\limits }\)
\(\newcommand {\cirfnint }{\mathop {\unicode {x2A10}}\limits }\)
\(\newcommand {\awint }{\mathop {\unicode {x2A11}}\limits }\)
\(\newcommand {\rppolint }{\mathop {\unicode {x2A12}}\limits }\)
\(\newcommand {\scpolint }{\mathop {\unicode {x2A13}}\limits }\)
\(\newcommand {\npolint }{\mathop {\unicode {x2A14}}\limits }\)
\(\newcommand {\pointint }{\mathop {\unicode {x2A15}}\limits }\)
\(\newcommand {\sqint }{\mathop {\unicode {x2A16}}\limits }\)
\(\newcommand {\intlarhk }{\mathop {\unicode {x2A17}}\limits }\)
\(\newcommand {\intx }{\mathop {\unicode {x2A18}}\limits }\)
\(\newcommand {\intcap }{\mathop {\unicode {x2A19}}\limits }\)
\(\newcommand {\intcup }{\mathop {\unicode {x2A1A}}\limits }\)
\(\newcommand {\upint }{\mathop {\unicode {x2A1B}}\limits }\)
\(\newcommand {\lowint }{\mathop {\unicode {x2A1C}}\limits }\)
\(\newcommand {\bigtriangleleft }{\mathop {\unicode {x2A1E}}\limits }\)
\(\newcommand {\zcmp }{\mathop {\unicode {x2A1F}}\limits }\)
\(\newcommand {\zpipe }{\mathop {\unicode {x2A20}}\limits }\)
\(\newcommand {\zproject }{\mathop {\unicode {x2A21}}\limits }\)
\(\newcommand {\biginterleave }{\mathop {\unicode {x2AFC}}\limits }\)
\(\newcommand {\bigtalloblong }{\mathop {\unicode {x2AFF}}\limits }\)
\(\newcommand {\arabicmaj }{\mathop {\unicode {x1EEF0}}\limits }\)
\(\newcommand {\arabichad }{\mathop {\unicode {x1EEF1}}\limits }\)
10.3 Finding Matches Quickly¶
The main advantage that LZ77 compression has over Huffman compression is that it is able to compress not just individual bytes but longer sequences of bytes. The downside of this approach is that it is more computationally intensive: Instead of simply translating each byte from the input into
a fixed bit sequence, LZ77 has to repeatedly scan the entire dictionary for the longest sequence of bytes that matches the current input. When the dictionary region is small, as in our examples so far, we can scan the dictionary and find the longest match using “brute force,”
by iterating over the entire dictionary and comparing each position to the current input.
To achieve good compression rates, however, the dictionary region must usually be significantly larger: Typical values are in the range of 10 000 to 1 000 000 bytes. For dictionaries of this size, finding the longest match using brute force becomes
impracticable. Much of the research literature on Lempel-Ziv compression is therefore concerned with different algorithms and data structures for efficiently finding the longest match. The idea behind most of these approaches is to build an auxiliary data structure that lets us quickly find positions in the
dictionary that match (or at least resemble) the bytes in the lookahead region. We will call this data structure the index, because it serves a similar purpose as the index in a book.
Indexing the Dictionary Region
The main purpose of the index is to find all positions in the dictionary that match a given string. One of the simplest data structure we can use for this purpose is a table that keeps track of where in the dictionary region each possible character or byte occurs. For example, assume we have advanced to the
following position in the input:
The dictionary contains the letter A at positions 1 and 4, the letter B at positions 2 and 3, the space character at position 5, and the letter D at position 6, so our index
table would look as follows:
| . |
| byte |
positions |
| A |
1, 4 |
| B |
2, 3 |
| D |
6 |
| ␣ |
5 |
When searching for the longest match of the current input ABBA, we can use the index to find all occurrences of the first letter A in the dictionary. In this case, A is found in two positions: The byte sequence
at position 1 matches all four letters of ABBA, whereas the one at position 4 matches only the first letter. Since the other positions in the dictionary do not even match the first letter, we don’t have to examine them at all.
We can extend this idea to longer matches by maintaining a table that contains the positions of all three-byte sequences in the dictionary:
| . |
| bytes |
positions |
| ABB |
1 |
| A␣D |
4 |
| BA␣ |
3 |
| BBA |
2 |
| DAB |
6 |
| ␣DA |
5 |
Notice that this table includes the 3-byte sequences starting at positions 5 and 6 even though they extend beyond the end of the dictionary region; this is useful for LZ77, which must be able to find matches that overlap the current input.
The benefit of indexing multiple bytes is that finding long matches becomes much faster. In this particular example, the three-byte sequence “ABB” occurs at only one position, so we find the match for the current input “ABBA” immediately. One
limitation of the extended index is that we lose the ability to find matches that are shorter than three bytes; such an index can therefore only be used if we aren’t searching for shorter matches anyway, that is, if the MinLength parameter is at least 3.
What is an appropriate data structure for storing such an index table? For an index that stores the positions of individual bytes, we can get by with a simple nested list:
List<List<Integer>> index;
The outer list consists of 256 entries, one for each possible byte, and the inner lists contains all locations of the corresponding byte. If we want to index the positions of byte sequences instead of individual bytes, such a simple data structure quickly becomes impractical, mainly because the necessary
size of the outer list increases rapidly. For example, when indexing three-byte sequences, the required size of index grows to \(256^3=16\,777\,216\) entries. Even worse, many of these entries will be empty since only a small fraction of all possible three-byte sequences is
likely to occur in real-world files. To avoid wasting memory, we therefore need to replace the outer array with data structure that doesn’t reserve memory for potentially millions of non-existent index entries.
One way to obtain such a data structure is to employ some form of hashing to redistribute the index entries over a smaller range of values. We can do this by defining a hash function \(h(\cdot )\) that maps each byte sequence to a small numeric value; the number \(h(x)\) is also referred to as the hash code of \(x\). For example, we can define a hash function for 3-byte sequences by simply computing the sum of the three individual bytes:
\(\seteqnumber{0}{10.}{0}\)
\begin{equation*}
h(b_1b_2b_3) = b_1+b_2+b_3
\end{equation*}
Let’s use this function to compute the hash codes of all 3-byte sequences in the same example as before:
The hash code for the first byte sequence ABB is \(h(\mathtt {ABB}) = 65 + 66 + 66 = 197\). (We use the ASCII table in Fig. 9.2 to translate letters to numeric codes.) The hash code for the second byte
sequence BBA is \(66 + 66 + 65 = 197\) as well, and for the remaining positions we obtain the following table:
| . |
| position |
bytes |
hash code |
| 1 |
ABB |
\(197\) |
| 2 |
BBA |
\(197\) |
| 3 |
BA␣ |
\(163\) |
| 4 |
A␣D |
\(165\) |
| 5 |
␣DA |
\(165\) |
| 6 |
DAB |
\(199\) |
In this case, the six byte sequences map to just four different hash codes; the code \(197\) occurs at positions 1 and 2, the code \(163\) at position \(3\), the code \(165\) at positions 4 and 5, and the code \(199\) at position 6:
| . |
| hash code |
positions |
| 163 |
3 |
| 165 |
4, 5 |
| 197 |
1, 2 |
| 199 |
6 |
The largest hash code that can occur with this particular hash function is \(256\cdot 3\), so we have managed to reduce the size of the index table from 16 million entries to just 768.
Searching the index works almost the same as before. To find the longest match of the current string ABBA, we compute the hash code of its first three characters, which is \(65+66+66=197\), and then consult the index to find all matching positions in the dictionary, in this
case positions 1 and 2. We then then iterate over all matching positions and compare them to ABBA to find the longest match. The byte sequence at position 1 is ABBA and therefore a match of length 4, whereas the one
at position 2 is BBA and therefore not a match.
As we just saw, it is possible for distinct byte sequences to have the same hash code, so an index implemented using hashing doesn’t contain guaranteed matches but only potential matches. When two distinct values have the same hash code, we say they collide in the hash table. The likelihood of collisions is determined by the hash function, specifically by the range of values it produces and how well it distributes its arguments over this range. The hash function we used until now — the sum of all
bytes — fares relatively poorly in both regards; we will discuss a better alternative shortly.
Implementing the Index
Let’s now turn to the implementation of the ideas we have discussed so far. The index maps integers (hash codes) to lists of integers (positions in the dictionary), so we could define it simply as a nested list of integers, as before
List<List<Integer>> index;
The code in Listing 10.9 uses a slightly different definition for index that replaces the inner List with a double-ended queue of type Deque. As we will explain in more detail below, this allows us to add new entries at one end of the queue and remove outdated entries from the other. The initIndex() method in the same listing initializes index by allocating an array of size INDEX_SIZE and populating each entry with an empty queue.
Listing 10.9ch10 / LZTokenizer
private final int INDEX_SIZE = 8191;
private List<Deque<Integer>> index;
private void initIndex() {
index = new ArrayList<>(INDEX_SIZE);
for (int i = 0; i < INDEX_SIZE; i++)
index.add(new ArrayDeque<>());
}
Next, we need to define a hash function that maps sequences of bytes to entries in index. Instead of simply computing the sum of all bytes, as we did previously, we use a function that spreads the resulting hash codes more evenly and over a
larger range. More specifically, we compute the hash code of a byte sequence \(b_0,\ldots ,b_k\) by combining their numerical values as follows:1
\(\seteqnumber{0}{10.}{0}\)
\begin{equation*}
h(b_0,\ldots ,b_k)= b_0 31^k + b_1 31^{k-1} +\cdots + b_{k-1} 31 + b_k = \sum _{i=0}^k b_i 31^{k-i}
\end{equation*}
We can obtain an equivalent expression that can be evaluated more efficiently by repeatedly factoring out the multiplier \(31\):
\(\seteqnumber{0}{10.}{0}\)
\begin{equation*}
h(b_0,\ldots ,b_k)=(\cdots (b_0\cdot 31+b_1)\cdot 31+\cdots )\cdot 31 + b_k
\end{equation*}
This factorization is known as Horner’s scheme and is widely used when computing sums of powers; we will meet it again in a different context in the Chapter 11.
The implementation of this hash function is straightforward. In our case, the hash code must only be computed for the byte sequence at the front of window. By default, the hashFront() method in Listing 10.10 computes the hash code of the next minLength
bytes in the lookahead region, but fewer bytes are used near the end of the input. The return statement in the last line reduces the computed hash code to a number between 0 and INDEX_SIZE-1. The floorMod() method is used instead
of the % operator to ensure that the result is always a nonnegative integer.
Listing 10.10ch10 / LZTokenizer
// Compute the hash code of the bytes at positions
// front, front+1, ...
private int hashFront() {
int length = Math.min(minLength, lookaheadSize);
int hash = 0;
for (int i = 0; i < length; i++) {
hash = hash * 31 + window.get(front + i);
}
return Math.floorMod(hash, INDEX_SIZE);
}
The index must be updated every time the input is advanced and new characters enter the dictionary region. The advanceInput() method shown in Listing 10.11 advances the input
by numBytes characters and uses a for-loop to update both index and window for each character read. First, we use hashFront() to retrieve the
index entry for the byte sequence at the front of the lookahead region; this is the byte sequence that is about to enter the dictionary. We then associate the current position with this hash code by adding front to positions. Since we add new positions using
addFirst(), the positions in each index entry are in reverse order: The first element is the latest matching position in the file and the last element the earliest. Afterward, advanceInput() reads the next character from the input, writes it to window, and increments front to shift the window one place to the right. Once we have reached the end of the input, we gradually decrease lookaheadSize.
Listing 10.11ch10 / LZTokenizer
// Advance the input and update 'window' and 'index'.
private void advanceInput(int numBytes) throws IOException {
assert lookaheadSize >= numBytes;
for (int i = 0; i < numBytes; i++) {
Deque<Integer> positions = index.get(hashFront());
positions.addFirst(front);
int nextByte = input.read();
if (nextByte != -1)
window.set(front + lookaheadSize, (byte) nextByte);
else
lookaheadSize--;
front++;
}
}
Next we need a method that queries the index to find all positions in the dictionary that potentially match the byte sequence at the front of the lookahead region. The simplest possible implementation of this operation simply returns the index entry for hashFront():
return index.get(hashFront())
The findMatches() method in Listing 10.12 additionally performs a bit of maintenance work and first removes all
positions from the entry that are no longer inside the dictionary region. Since the positions in each queue are stored in decreasing order, we can simply drop all positions from the end of the queue that are smaller than dictStart, the current start of the dictionary.
Listing 10.12ch10 / LZTokenizer
// Find the positions of potential matches in the dictionary.
private Iterable<Integer> findMatches() {
Deque<Integer> positions = index.get(hashFront());
int dictStart = front - maxOffset;
while (!positions.isEmpty() && positions.getLast() < dictStart)
positions.removeLast();
return positions;
}