All rights reserved. Copyright (C) 1996-2005 by NARITA Tomio
Last modified at Jan.16th,2004.
LV Homepage
-
-
lv - a Powerful Multilingual File Viewer / Grep
The latest version is ver 4.51:
Download
Table of Contents
-
- Copyright
- Feature
- Download lv
- Installation
- Usage
- Limitations
- Coding systems
- Annotation about encoding/decoding scheme
- Auto selection of a coding system
- Extension for text decoration
- Customization
- Bug report
- Release note
- Acknowledgement
- Reference
Copyright
-
-
All rights reserved. Copyright (C) 1996-2005 by NARITA Tomio.
This program is free software; you can redistribute it and/or modify
it under the terms of the GNU General Public License as published by
the Free Software Foundation; either version 2 of the License, or
(at your option) any later version.
This program is distributed in the hope that it will be useful,
but WITHOUT ANY WARRANTY; without even the implied warranty of
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
GNU General Public License for more details.
You should have received a copy of the GNU General Public License
along with this program; if not, write to the Free Software
Foundation, Inc., 59 Temple Place, Suite 330, Boston, MA 02111-1307 USA
See also GNU General Public License Version 2.
Feature
-
Multilingual file viewer
lv is a powerful multilingual file viewer.
Apparently, lv looks like less (1),
a representative file viewer on UNIX as you know,
so UNIX people (and less people on other OSs)
don't have to learn a burdensome new interface.
lv can be used on MSDOS ANSI terminals and almost all UNIX platforms.
lv is a currently growing software,
so your feedback is welcome
and helpful for us to refine the future lv.
-
Multiple coding systems
lv can decode and encode multilingual streams
through many coding systems, for example,
ISO 2022 based coding systems such as iso-2022-jp,
and EUC (Extended Unix Code) like euc-japan.
Furthermore,
localized coding systems
such as shift-jis, big5 and HZ are also supported.
lv can be used not only as a file viewer
but also as a coding-system translation filter
like nkf (1) and tcs (1).
-
Multilingual regular expressions / Multilingual grep
lv can recognize multi-bytes patterns as regular expressions,
and lv also provides multilingual grep (1) functionality
by giving it another name, lgrep.
Pattern matching is conducted in the charset level,
so an EUC fragment, for example,
can be found in the ISO 2022 tailored streams, of course.
-
Supporting the Unicode standard
lv provides Unicode facilities
which enables you to handle Unicode streams encoded in UTF-7 or UTF-8,
and lv can also convert their code-points
between Unicode and other charsets.
So you can display Unicode or foreign texts on your terminal,
using the code conversion function
to your favorite charsets via Unicode.
(However, MSDOS version of lv has none of the Unicode facility.)
-
ANSI escape sequence through
lv can recognize ANSI escape sequences for text decoration.
So you can look ANSI-decorated streams
such as colored source codes generated by another software
just like intended image on ANSI terminals.
-
Completely original
lv is a completely original software
including no code drawn from less and grep
and other programs at all.
Sample Images
Download lv
-
-
You can download lv archive.
Changes between older versions are described in
release note
(in Japanese).
Installation
Standard installation:
- Expand lv archive, using gunzip/tar.
- Change your working directory to ``(extracted sub directory)/build''.
- Execute ``../src/configure'' to configure compiler flags.
- Launch ``make''.
- Then, launch ``make install'' as root.
MSDOS installation:
Before making lv,
you need to install
LSI C-86 Compiler
(limited and freeware version of LSI C-86 for sample usage).
- Expand lv archive, using gunzip/tar.
- Change your working directory to ``(extracted sub directory)/src''.
- Launch ``make -f Makefile.dos''.
- Copy ``lv.hlp'', brief help description, to the same directory
as lv.exe settled.
MSDOS version of lv directly outputs ANSI escape sequences
without regard to termcap and terminfo.
Perhaps you need an ANSI escape sequence driver named ``ANSI.SYS''
(or more sophisticated one) on MSDOS
including DOS prompt on MS-Windoze.
Since Windoze-NT does not seem to prepare such drivers
for DOS prompt in default,
please look into the driver configuration
when lv fails to handle the terminal capability correctly.
Usage
-
How to launch lv?
When you just wish to display a file on a terminal,
please launch lv from command line like this:
-
-
% lv [options] files ...
Or, using redirect or pipe-line:
-
-
% another_command | lv [options]
% lv [options] < file
Compressed files that have suffix ``gz'', ``z'', or ``GZ'', ``Z'' are
extracted by lv using zcat (1),
and ``bz2'' or ``BZ2'' with bzcat (1).
Please install zcat and bzcat that can expand all of them.
In case that standard output is not connected to an ordinal terminal
but to redirect or pipe-line,
lv works as a coding-system or code-points conversion filter
like nkf (1) and tcs (1).
lv also works like grep (1)
by giving it another name, lgrep.
Please install symbolic (or hard) link
whose name is lgrep to lv (1).
Or, lgrep functionality is also turned on the option '-g'.
lgrep is used like below:
-
-
% lgrep [options] grep_pattern files ...
% another_command | lgrep [options] grep_pattern
% lgrep [options] grep_pattern < file
The coding-system of grep_pattern can be specified
as ``keyboard coding system'' (see below).
-
Command line options
- -A<coding-system>
- Set all coding systems to coding-system.
- -I<coding-system>
- Set input coding system to coding-system.
- -K<coding-system>
- Set keyboard coding system to coding-system.
If it is not set, output coding system will be applied to it.
- -O<coding-system>
- Set output coding system to coding-system.
- -P<coding-system>
- Set pathname coding system to coding-system.
- -D<coding-system>
- Set default EUC coding system to coding-system.
-
coding-system
-
- a: auto-select
Its entity is iso-2022-kr
until an 8bit code is found.
- c: iso-2022-cn
- j: iso-2022-jp
- k: iso-2022-kr
- e: Extended Unix Code
- ec: euc-china
- ej: euc-japan
- ek: euc-korea
- et: euc-taiwan
- u: UCS transformation format
- l: iso-8859-1..9
- l1..9: iso-8859-1..9
- l0: iso-8859-10
- lb,ld,le,lf,lg: iso-8859-11,13,14,15,16
- s: shift-jis
- b: big5
- h: HZ
- r: raw mode
No decoding and encoding are performed.
Coding-system translations / Code-points conversions:
iso-2022-cn, -jp, -kr can be converted into euc-china or -taiwan,
euc-japan, euc-korea, respectively (and vice versa).
shift-jis uses the same internal code-points
as iso-2022-jp and euc-japan.
Since big5 characters can be converted into CNS 11643-1992
with negligible incompleteness,
big5 streams can be translated into iso-2022-cn or euc-taiwan
(and vice versa) with code-points conversion.
Note that the iso-2022-cn referred here is not GB sequence,
only just CNS one.
You should remember that lv cannot translate big5 into GB directly.
The search function of lv may not work correctly when lv additionally
performs ``code-points'' conversion
(not ``coding-system'' translation),
because visible code and internal code are different from each other.
lv will try to avoid this problem with
converting charsets of search patterns automatically,
but this function is not always perfect.
- -W<number>
- Screen width
- -H<number>
- Screen height
- -E'<editor>'
- Editor name (default 'vi -c %d')
``%d'' means the line number of current position in a file.
- -q
- Assert there is delete/insert-lines control
Please set this option on a MSDOS ANSI terminal
that has capability to delete and/or insert lines.
As to termcap and terminfo version,
it will be set automatically.
- -Ss<seq>
- Set ANSI Standout sequence to (default "7")
- -Sr<seq>
- Set ANSI Reverse sequence to (default "7")
- -Sb<seq>
- Set ANSI Blink sequence to (default "5")
- -Su<seq>
- Set ANSI Underline sequence to (default "4")
- -Sh<seq>
- Set ANSI Highlight sequence to (default "1")
These sequences are inserted
between ``ESC ['' and ``m''
to construct full ANSI escape sequences.
- -T<number>
-
Set Threshold-code which divides Unicode code-points in
two regions. Characters belonging to the lower region are
assumed to have a width of one, and the higher characters
are equated to a width of two. (Default: 12288, = 0x3000)
- -m
-
Force Unicode code-points which have the same glyphs as
iso-8859-* to be Mapped to iso-8859-* in a conversion from
Unicode to another character set which also has the
corresponding code-points, in particular, Asian charsets.
- -a
- Adjust character set for search pattern (default)
- -c
- Allow ANSI escape sequences for text decoration (Color)
- -d, -i
- Make regexp-searches ignore case (case folD search)
(default)
- -f
- Substitute Fixed strings for regular expressions
- -k
- Convert X0201 Katakana to X0208
- -l
- Allow physical lines of each logical line printed
on the screen to be concatenated for cut and paste
after screen refresh
- -s
- Force old pages to be swept out from the screen Smoothly
- -u
- Unify several character sets, eg. JIS X0208 and C6226.
In addition, lv equates ISO 646 variants,
eg. JIS X0201-Roman,
and unknown charsets with ASCII.
- -g
- Turn on lgrep mode.
- -n
- Prefix each line of output with the line number within its input file on lgrep.
- -v
- Invert the sense of matching on lgrep.
- -z
- Enable HZ auto-detection (also enabled by run-time C-t).
- -+
- Clear all options
You can also turn OFF specified options,
using ``+<option>'' like +c, +d, ... +z.
- -
- Treat the following arguments as filenames
- -V
- Show lv version
- -h
- Show this help
-
Configuration
Options can be described in the configuration file ``.lv''
(``_lv'' on MSDOS) located at you home directory. If and only if you
use MSDOS, you can locate ``_lv'' at current working directory.
They can be also described in the environment variable LV.
Every configuration will be overloaded in the following order if there is.
Command line options are always read finally.
- .lv located at your home directory
- (_lv located at current working directory: MSDOS only)
- Environment variable LV
- Command line options
Examples:
- MSDOS (Input is shift-jis, Screen height is 25 lines, Highlight seq is "1;45", Underline seq is "1")
set LV=-Is -H25 -Sh1;45 -Su1
- UNIX csh (Input is HZ-enabled auto-select, Output and Keyboard is both iso-2022-cn)
setenv LV '-z -Oc -Dec'
-
Run-time commands
- 0-9:
- Argument
- g, <:
- Jump to the line number (default: top of the file)
- G, >:
- Jump to the line number (default: bottom of the file)
- p:
- Jump to the percentage position in line numbers (0-100)
- b, C-b:
- Previous page
- u, C-u:
- Previous half page
- k, w, C-k, y, C-y, C-p:
- Previous line
- j, C-j, e, C-e, C-n, CR:
- Next line
- d, C-d:
- Next half page
- f, C-f, C-v, SP:
- Next page
- F:
- Jump to the end of file, and wait for a data to be
appended to the file until interrupted.
- /<string>:
- Find a string in the forward direction (regular expression)
- ?<string>:
- Find a string in the backward direction (regular expression)
- n:
- Repeat previous search in the forward direction
- N:
- Repeat previous search in the backward direction (not REVERSE)
- C-l:
- Redisplay all lines
- r, C-r:
- Refresh screen and memory
- R:
- Reload the current file
- :n:
- Examine the next file
- :p:
- Examine the previous file
- t:
- Toggle input coding systems
- T:
- Toggle input coding systems reversely
- C-t:
- Toggle HZ decoding mode
- v:
- Launch the editor defined by option -E
- C-g, =:
- Show file information (filename, position, coding system)
- V:
- Show LV version
- C-z:
- Suspend (call SHELL or ``command.com'' under MSDOS)
- q, Q:
- Quit
- UP/DOWN:
- Previous/Next line
- LEFT/RIGHT:
- Previous/Next half page
- PageUp/PageDown:
- Previous/Next page
-
How to input search strings?
You can input a string which consists of multi-bytes characters
and search the string as a regular expression.
lv's regular expression is similar to Mule's one.
The following keys have special meanings in the keyboard input:
- C-m, Enter
- Enter the current string
- C-h, BS, DEL
- Delete one character (backspace)
- C-u
- Cancel the current string and try again
- C-p
- Restore a few old strings incrementally (history)
- C-g
- Quit
-
Regular expressions
- `. (period)'
matches any single character.
For example,
``a.b'' matches any three-character string which begins with
`a' and ends with `b'.
- `*'
constructs repetition of an expression more than 0 times.
For example,
``ab*'' matches `a', `ab' `abb', etc.
- `+'
constructs repetition of an expression more than once.
For example,
``ab+'' matches `ab', `abb', but not `a'.
- `?'
matches the preceding expression either once or not at all.
For example,
``ca?r'' matches `car' or `cr'; nothing else.
- `[ ... ]'
makes a character set.
For example,
``[ab]+'' matches any string composed of just `a's and `b's.
You can also include character ranges in a character set,
by writing two characters with a `-' between them.
For example,
``[a-z]'' matches any lower-case letter.
If the characters implies a multi-bytes charset,
lv makes a multi-bytes range,
ordering code-points as unsigned integer.
Mutually overlapping ranges (or charset) are not guaranteed.
- `[^ ... ]'
makes a complemented character set.
For example,
``[^a-z0-9A-Z]'' matches all characters
*except* letters and digits.
- `^'
matches the empty string at the beginning of a line.
- `$'
is similar to `^' but matches only at the end of a line.
- `\'
quotes the special characters.
- `\1'
matches characters each of which has a width of 1 column.
- `\2'
matches characters each of which has a width of 2 columns.
- `\|'
specifies an alternative.
For example,
``foo\|bar'' matches either `foo' or `bar' but no other string.
- `\( ... \)'
\(, \) is a grouping construct.
For example,
``ba\(na\)*'' matches `ba', `bana', `banana', etc.
Limitations
-
Up to 8192 bytes per a logical line
lv manages file location pointers logically,
separating LOGICAL lines by LF (line feed) or CR (carriage return),
or CR/LF.
The length of a logical line is limited up to 8192 bytes.
And lv insert a LF forcibly when a line has a length over 8192 bytes.
Note that all of CRs or CR/LF are replaced with single LF on UNIX
during decoding.
As to MSDOS,
CRs are inserted before every LFs without thinking.
-
Physical lines per a logical line
A logical line is divided into PHYSICAL lines
to fall into the screen width.
lv limits physical lines up to "characters / 16" lines length
per a logical line for management of them.
Note that when a logical line has more lines,
the rest of the limit are truncated and not displayed at all.
-
Limitation of encoding space
Encoding space is limited upto "characters * 4" bytes length
for each decoded string.
Even if encoded string would be longer than that,
the encoding process is dropped at the limit.
-
Limitation of the number of logical lines
The number of logical lines is also limited.
Currently,
lv can handle up to about 2 Giga lines on UNIX
(65000 lines on MSDOS).
Note that lines which exceed this limitation cannot be displayed at all.
Coding systems
-
ISO 2022 based coding systems
lv handles ISO 2022 based coding systems as
they are stateless on the logical line level.
So you have to specify a coding system before decoding,
and lv maybe adds redundant codes during encoding.
- iso-2022-cn
RFC 1922 tailored coding system.
| | G0 | G1 | G2 | G3
|
| Designation | ASCII | GB 2312-80, CNS 11643-1992 Plane 1, ISO-IR-165 | CNS 11643-1992 Plane 2 | CNS 11643-1992 Plane 3..7
|
- iso-2022-jp
RFC 1468 and 1554 tailored coding system.
All 94charsets use G0, and all 96charsets use G2 with single shift
inside lv.
- iso-2022-kr
RFC 1557 tailored coding system.
All charsets except ASCII use only G1 with locking shift
inside lv.
-
Extended Unix Code
lv can decode mixture texts of euc-* and iso-2022-*,
when you select euc-* as the input coding system.
- euc-china
| | G0 | G1 | G2 | G3
|
| Designation | ASCII | GB 2312-80 | not used | not used
|
- euc-japan
| | G0 | G1 | G2 | G3
|
| Designation | ASCII | JIS X 0208 | JIS X 0201 Katakana | JIS X 0212
|
- euc-korea
| | G0 | G1 | G2 | G3
|
| Designation | ASCII | KS C 5601-1987 | not used | not used
|
- euc-taiwan
| | G0 | G1 | G2 | G3
|
| Designation | ASCII | CNS 11643 Plane 1 | CNS 11643 Plane 2-7 | not used
|
-
UCS transformation format
lv can convert character codesets
between Unicode and the following charsets:
GB 2312-80, JIS X 0208, JIS X 0212, KSC 5601-1987,
Big Five, CNS 11643-1992 Plane 1-2,
and ISO 8859-1..16.
Currently lv's mapping table is based on Unicode 1.1.
| Encoding | Charset used for mapping from Unicode
|
| iso-2022-cn | GB 2312-80 (primary), CNS 11643-1992 (secondary), (ISO 8859-*)
|
| iso-2022-jp | JIS X0208, JIS X0212, JIS X0201, (ISO 8859-*)
|
| iso-2022-kr | KSC 5601-1987, (ISO 8859-*)
|
| euc-china | GB 2312-80
|
| euc-japan | JIS X0208, JIS X0212, JIS X0201
|
| euc-korea | KSC 5601-1987
|
| euc-taiwan | CNS 11643-1992 Plane 1-2
|
| shift-jis | JIS X0208, JIS X0201
|
| big5 | Big Five
|
When you output Unicode CJK unified ideographs through iso-2022-cn,
GB 2312-80 is used primarily,
and the rest which are not included in GB
are mapped into CNS 11643-1992.
-
Other coding systems
- Limitations
- Coding systems
- Annotation about encoding/decoding scheme
- Auto selection of a coding system
- Extension for text decoration
- Customization
- Bug report
- Release note
- Acknowledgement
- Reference
Copyright
-
-
All rights reserved. Copyright (C) 1996-2005 by NARITA Tomio.
This program is free software; you can redistribute it and/or modify
it under the terms of the GNU General Public License as published by
the Free Software Foundation; either version 2 of the License, or
(at your option) any later version.
This program is distributed in the hope that it will be useful,
but WITHOUT ANY WARRANTY; without even the implied warranty of
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
GNU General Public License for more details.
You should have received a copy of the GNU General Public License
along with this program; if not, write to the Free Software
Foundation, Inc., 59 Temple Place, Suite 330, Boston, MA 02111-1307 USA
See also GNU General Public License Version 2.
Feature
-
Multilingual file viewer
lv is a powerful multilingual file viewer.
Apparently, lv looks like less (1),
a representative file viewer on UNIX as you know,
so UNIX people (and less people on other OSs)
don't have to learn a burdensome new interface.
lv can be used on MSDOS ANSI terminals and almost all UNIX platforms.
lv is a currently growing software,
so your feedback is welcome
and helpful for us to refine the future lv.
-
Multiple coding systems
lv can decode and encode multilingual streams
through many coding systems, for example,
ISO 2022 based coding systems such as iso-2022-jp,
and EUC (Extended Unix Code) like euc-japan.
Furthermore,
localized coding systems
such as shift-jis, big5 and HZ are also supported.
lv can be used not only as a file viewer
but also as a coding-system translation filter
like nkf (1) and tcs (1).
-
Multilingual regular expressions / Multilingual grep
lv can recognize multi-bytes patterns as regular expressions,
and lv also provides multilingual grep (1) functionality
by giving it another name, lgrep.
Pattern matching is conducted in the charset level,
so an EUC fragment, for example,
can be found in the ISO 2022 tailored streams, of course.
-
Supporting the Unicode standard
lv provides Unicode facilities
which enables you to handle Unicode streams encoded in UTF-7 or UTF-8,
and lv can also convert their code-points
between Unicode and other charsets.
So you can display Unicode or foreign texts on your terminal,
using the code conversion function
to your favorite charsets via Unicode.
(However, MSDOS version of lv has none of the Unicode facility.)
-
ANSI escape sequence through
lv can recognize ANSI escape sequences for text decoration.
So you can look ANSI-decorated streams
such as colored source codes generated by another software
just like intended image on ANSI terminals.
-
Completely original
lv is a completely original software
including no code drawn from less and grep
and other programs at all.
Sample Images
Download lv
-
-
You can download lv archive.
Changes between older versions are described in
release note
(in Japanese).
Installation
Standard installation:
- Expand lv archive, using gunzip/tar.
- Change your working directory to ``(extracted sub directory)/build''.
- Execute ``../src/configure'' to configure compiler flags.
- Launch ``make''.
- Then, launch ``make install'' as root.
MSDOS installation:
Before making lv,
you need to install
LSI C-86 Compiler
(limited and freeware version of LSI C-86 for sample usage).
- Expand lv archive, using gunzip/tar.
- Change your working directory to ``(extracted sub directory)/src''.
- Launch ``make -f Makefile.dos''.
- Copy ``lv.hlp'', brief help description, to the same directory
as lv.exe settled.
MSDOS version of lv directly outputs ANSI escape sequences
without regard to termcap and terminfo.
Perhaps you need an ANSI escape sequence driver named ``ANSI.SYS''
(or more sophisticated one) on MSDOS
including DOS prompt on MS-Windoze.
Since Windoze-NT does not seem to prepare such drivers
for DOS prompt in default,
please look into the driver configuration
when lv fails to handle the terminal capability correctly.
Usage
-
How to launch lv?
When you just wish to display a file on a terminal,
please launch lv from command line like this:
-
-
% lv [options] files ...
Or, using redirect or pipe-line:
-
-
% another_command | lv [options]
% lv [options] < file
Compressed files that have suffix ``gz'', ``z'', or ``GZ'', ``Z'' are
extracted by lv using zcat (1),
and ``bz2'' or ``BZ2'' with bzcat (1).
Please install zcat and bzcat that can expand all of them.
In case that standard output is not connected to an ordinal terminal
but to redirect or pipe-line,
lv works as a coding-system or code-points conversion filter
like nkf (1) and tcs (1).
lv also works like grep (1)
by giving it another name, lgrep.
Please install symbolic (or hard) link
whose name is lgrep to lv (1).
Or, lgrep functionality is also turned on the option '-g'.
lgrep is used like below:
-
-
% lgrep [options] grep_pattern files ...
% another_command | lgrep [options] grep_pattern
% lgrep [options] grep_pattern < file
The coding-system of grep_pattern can be specified
as ``keyboard coding system'' (see below).
-
Command line options
- -A<