-
DBFCSV: Conversion between DBF and CSV files
Direct access to online help: DBFCSV
Access the application from the menu:
"File | Import | Vectors and/or databases | DBF, Access, etc -> PNT, ARC, POL"
"File | Export | DBF -> CSV"
Presentation and options
This application converts files from DBF format to CSV format and vice versa. TXT files are also supported as a tabular data format, both for reading and writing. The available options are:
From DBF to CSV (option 1):
Converts files from the DBF format (dBASE tables in classic DBF format or in extended DBF) to the CSV format (export of Excel tables)
From CSV to DBF (option 2):
Converts files from the CSV format (export of Excel tables) to DBF format (dBASE tables in classic DBF format or in extended DBF).
In the DBF to CSV/TXT conversion, the data from the original table are transformed into a text file. In this file, the text appears on the same line, and the different columns or fields are distinguished by a separator character, which can be selected. In many European countries, the preferred separator is the semicolon (;), while the comma (,) is commonly used in the United States and many other countries. This separator can also be a tab character; in this case, the text TAB must be entered as the list separator. The tab character is an excellent (recommended) choice because reading CSV files in other programs (such as MS Excel itself) is much less problematic when fields contain commas, semicolons, quotation marks, apostrophes, etc. The special words SPA, to indicate that a space will be used as the list separator, and CAP, to indicate that there is no list separator, are also accepted.
The first row of the generated CSV will contain the field names, which will become the columns of the CSV file. The subsequent rows will contain the values of each record, separated by the selected delimiter.
In the reverse conversion, from CSV/TXT to DBF, you must select the separator used in the CSV to generate the different columns (if in doubt, you can open the file to check which separator is being used through the icon
or open it with any text editor; keep in mind that, if quotation marks are present, they normally delimit text that will be placed in the same field, or, when doubled (""), they indicate a single quotation mark that should be retained as a text character, typically to represent the arc-second symbol, or seconds as 1/60 of a minute).
The first row of the CSV may contain the column names (when First line with header is selected); in this case, these names can be used as the field names of the generated DBF. In this first row, the separator between the field names is the same as in the remaining rows that will become records. If the CSV file contains blank lines at the end, particularly due to a final line break, these lines are not added to the DBF as empty records. If these records without information are desired, as many separators as there are fields in the line must be written.
It should be noted that, in the case of CSV files, the program analyzes the file beforehand to determine the character set and determine whether the file is ANSI, OEM, or UTF-8. The output DBF table is written using the character encoding specified by the JocCaracDBFPerDefecte= key in MiraMon.par, which can be modified using a plain text editor or from the "Help | Configure parameters" menu, in the "Zoom and General aspects" tab.
In MiraMon, the main tables associated with files that are layers containing geographic or geometric content (graphic layers, etc.) are in DBF or extended DBF format, and always have a first column used to store what is known as the graphic identifier (ID_GRAFIC). This is used to provide each graphic object with a series of geometric-topological or thematic attributes. If the CSV does not have a first column of graphic identifiers, one can be generated by activating the "Add graphic identifier column" checkbox.
In addition, in MiraMon, DBF tables can exceed the limitations of the dBASE IV format (known as classic DBF; see the document dedicated to extended DBF -available in Catalan-), but this is not the case in other cases, such as tables corresponding to layers in Shape format. Therefore, the application provides the option to restrict the conversion from CSV to DBF to the limits of the classic DBF format (not extended); note that in this case, as expected, fields or parts of field contents may be lost, field names may need to be simplified, etc, because this information cannot be accommodated into a classic DBF.
It is also possible to specify which character is used as the text qualifier. In a CSV file, the text qualifier is used to delimit the beginning and end of text that is considered part of a column (which will become a field in the DBF). In other words, it indicates that everything between two qualifiers (for example, quotation marks or apostrophes) is part of the same field, even if it contains the character used to separate columns. This allows, for example, a text field to contain a comma when a comma has been specified as the separator, provided that the entire field content is enclosed in quotation marks, without the comma being interpreted as a column separator. For this purpose, the field content is enclosed between two characters known as text qualifiers.
MiraMon allows to select from 5 possible qualifiers:
- Apostrophe (')
- Quotation marks ("): this is the most common qualifier
- Apostrophe (except if the separator is TAB): when the column separator is the tab character, allows apostrophes that may appear within the text to be treated as normal characters rather than text qualifiers
- Quotation marks (except if the separator is TAB): similarly to the previous case, when quotation marks are found within the text. This is the default value
- None: no character is considered to act as a text delimiter
For example, in a record containing three fields: a numeric field (1), a text field (Pinus, Abies, etc), and another numeric field (28), with quotation marks selected as the qualifier:
1,"Pinus, Abies, etc",28
the commas within the text "Pinus, Abies, etc" are not considered separators, since the text has been specified as being delimited by the qualifier, which in this case is the quotation mark.
Finally, MiraMon also allows you to specify how a double qualifier should be interpreted. When a CSV file contains a double qualifier (that is, two consecutive qualifiers: ""), it may refer to a field with no content, but it may also refer to the " character itself, as when a field contains coordinates in degrees-minutes-seconds and the seconds value is followed by the " symbol. In these cases, the CSV file may have been written using a double qualifier "" to prevent the seconds symbol from being interpreted as the beginning of new text marked by a qualifier. This option allows you to specify how the presence of a double qualifier should be interpreted:
- Normal character: if the qualifier is, for example, QUOTATION MARKS ("), and the program finds "" in the CSV file, it will convert it to a single quotation mark (") in the DBF field. This is useful for supporting arc seconds in a string while also allowing strings with quotation marks as qualifiers. This is the default value
- Qualifier: the program will treat the double qualifier as a qualifier and, therefore, "" defines an empty string (an empty field in the DBF).
The application also supports reading and writing metadata files in CSVW (JSON) format to document column properties, such as field name, data type, maximum width, units (if specified), etc. When converting a CSV/TXT file to DBF, after selecting the CSV/TXT file, the application automatically attempts to locate the corresponding CSVW file and, if found, assigns it to the corresponding resource and updates the options available in the dialog box (list separator, presence of a header, etc.). If no CSVW file is found, or the user does not specify one, the application automatically attempts to determine the list separator and whether the CSV/TXT file contains a header.

Dialog box of the application

Examples
The following example shows the conversion of a DBF file containing information about a country's monumental trees.
If the semicolon (;) is used as the separator, the header and the first row of data from the original DBF table containing the monumental trees are written as follows:
ID_GRAFIC;MENA_DECLA;NOM_DECLAR;ESPECIE;INE;TERME_MUNI;COMARCA;MATRICULA;ESTAT_DETA;ESTAT_RESU;URL;COORX;COORY;OBJECTID
0;AM;Pi Vell de l'Arp II;Pinus uncinata;25909;Vansa i Fórnols, La;Alt Urgell, l';AM 04.909.01b;N;Viu;http://mediambient.gencat.cat/ca/05_ambits_dactuacio/patrimoni_natural/arbres-monumentals/am_arbres_monumentals_fitxes/alt-urgell-6/pins-vells-de-larp-l-ll-lll/;377432.00;4673067.00;26
|
| Metadata file in CSVW (JSON) format |

Syntax
Syntax:
- DBFCSV 1 FitxerOrigenDBF FitxerDestiCSV [/SEPARA] /CSVWOut
- DBFCSV 2 FieldIdGrafic Header DBFClassic SourceCSVFile TargetDBFFile [/SEPARA] [/DOBLE_QUALIF] [/QUALIF_TEXT] [/CSVW] [/GENERA_UTF8]
Options:
- 1:
Conversion from DBF to CSV.
- 2:
Conversion from CSV to DBF.
Parameters:
- FitxerOrigenDBF
(Input file -
Input parameter): Name of the DBF file to be converted.
- FitxerDestiCSV
(Output file -
Output parameter): Name of the CSV output file.
- FieldIdGrafic
(Field of the graphic identifier -
Input parameter): It has a value of 0 if a column with the field that contains the graphic identifier is desired to de added, if it exists; otherwise the value is 1.
- Header
(Header of the first row. -
Input parameter): It has a value of 1 if the first row of the CSV file is wanted to be used as the header of the DBF (which contains the names of the fields of the DBF).
- DBFClassic
(Restrict to classic DBF -
Input parameter): It has a value of 1 if the conversion of CSV to a non-extended DBF (classic) is restricted, otherwise the value is 0. In this case, the DBF will be correctly marked (if it is extended it can be seen in "Information | Table information" when opening the table in MiraDades).
- SourceCSVFile
(Input file -
Input parameter): Name of the CSV file to be converted.
- TargetDBFFile
(Output file -
Output parameter): Name of the DBF output file.
Modifiers:
/CSVWOut=
(Output CSVW file)
CSVW file, where the metadata of the tabular data content in the CSV file is stored. (Output parameter) /SEPARA= (Separator) When converting from DBF to CSV, the separator character for the CSV columns corresponding to the fields in the DBF table can be selected. The preferred separator in many European countries is the semicolon (;), while the comma (,) is preferred in the United States and many other countries. This modifier also accepts the special keyword TAB to indicate that the tab character (binary value 9 in the CSV file) will be used as the list separator. The tab character is an excellent and recommended choice because reading CSV files in other applications, such as MS Excel, is much less problematic when fields contain commas, semicolons, quotation marks, apostrophes, etc. The special keyword SPA is also accepted to indicate that a space will be used as the list separator, and CAP to indicate that no list separator is used. For the reverse conversion, from CSV to DBF, the separator used in the CSV file to generate the different columns must be selected. In case of doubt, the file can be opened in a text editor to determine which separator is being used. It should be kept in mind that quotation marks, when present, normally delimit text that will be stored in a single field. When doubled (""), they may instead be used to indicate a single quotation mark that is to be included as a text character, commonly to represent the arcsecond symbol or seconds as 1/60 of a minute. The default value is the semicolon when reading CSV files, whereas the tab character is the default when writing them. (Input parameter) /DOBLE_QUALIF=
(Double qualifier)
When a double qualifier is encountered in a CSV file (that is, two consecutive qualifiers: "") [see the meaning of qualifier in the explanation of the /QUALIF_TEXT modifier], it may refer to a field with no content, but it may also refer to the " character itself, as when a field contains coordinates in degrees-minutes-seconds and the seconds value is followed by the " symbol. In such cases, when the CSV file is written, a double qualifier "" may have been used to prevent the seconds symbol from being interpreted as the beginning of new text delimited by a qualifier. The /DOBLE_QUALIF modifier specifies how the presence of a double qualifier is to be interpreted: if CARACTER_NORMAL is specified and the qualifier is, for example, COMETES ("), and the program finds "" in the CSV file, it is interpreted as a single quotation mark (") in the DBF field. This is useful for supporting arcseconds within a string while also allowing strings to use quotation marks as qualifiers. If QUALIF is specified instead, the double qualifier is interpreted as having the function of a qualifier and, therefore, "" defines an empty string (an empty field in the DBF). CARACTER_NORMAL is the default value. (Input parameter) /QUALIF_TEXT=
(Text qualifier)
In a CSV file, the text qualifier is used to delimit the beginning and end of text that is to be considered part of a column (which will ultimately become a field in the DBF file). This makes it possible, for example, for a text field to contain a comma when /SEPARA has been set to a comma, provided that the entire field is enclosed in quotation marks used as text qualifiers. For example, a record containing a numeric field, a text field, and another numeric field may be represented as follows: 1,"Pinus, Abies, etc",28. The most common text qualifier is a quotation mark ("), although no qualifier may be defined, or an apostrophe (') or another character may be used instead. The /QUALIF_TEXT modifier specifies which character acts as the text qualifier: COMETES ("), APOSTROF ('), COMETES_EXCEPTE_SI_TAB (" except when /SEPARA=TAB has been specified, in which case any quotation mark found when the separator is a tab character is treated as an ordinary character), or APOSTROF_EXCEPTE_SI_TAB (' except when /SEPARA=TAB has been specified, in which case any apostrophe found when the separator is a tab character is treated as an ordinary character). If CAP is specified, no character is considered to act as a text delimiter. The default value is COMETES_EXCEPTE_SI_TAB because, when the list separator (/SEPARA) is a tab character (TAB), a string qualifier based on quotation marks or apostrophes is not required, and any quotation marks encountered are treated as ordinary characters. (Input parameter) /CSVW=
(CSVW file)
CSVW file, where the metadata of the tabular data content in the CSV file is read. (Input parameter) /GENERA_UTF8
(Generate UTF8)
Allows generating the DBF file in UTF-8 character encoding. (Input parameter)
