<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>http://lql.uni-trier.de/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=RVulanovic</id>
	<title>Laws in Quantitative Linguistics - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="http://lql.uni-trier.de/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=RVulanovic"/>
	<link rel="alternate" type="text/html" href="http://lql.uni-trier.de/index.php/Special:Contributions/RVulanovic"/>
	<updated>2026-09-18T20:09:38Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.31.0</generator>
	<entry>
		<id>http://lql.uni-trier.de/index.php?title=Complexity_of_syntactic_constructions&amp;diff=1860</id>
		<title>Complexity of syntactic constructions</title>
		<link rel="alternate" type="text/html" href="http://lql.uni-trier.de/index.php?title=Complexity_of_syntactic_constructions&amp;diff=1860"/>
		<updated>2006-12-11T17:26:17Z</updated>

		<summary type="html">&lt;p&gt;RVulanovic: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;'''1. Problem and history'''&lt;br /&gt;
&lt;br /&gt;
The complexity of a syntactic construct is measured in terms of the number of its immediate constituents. The partitioning in immediate constituents can be performed on the basis of any grammar.&lt;br /&gt;
The first model seems to be that of Köhler and Altmann (2000).&lt;br /&gt;
&lt;br /&gt;
'''2. Hypothesis'''&lt;br /&gt;
&lt;br /&gt;
''The complexity of syntactic constructions follows the hyper-Pascal distribution''.&lt;br /&gt;
&lt;br /&gt;
'''3. Derivation'''&lt;br /&gt;
&lt;br /&gt;
The complexity depends on following quantities (Köhler, Altmann 2000:192):&lt;br /&gt;
&lt;br /&gt;
minX – the requirement of minimization of the complexity of a syntactic construction in order to decrease memory effort in processing the construction;&lt;br /&gt;
&lt;br /&gt;
maxH – the requirement of maximazing compactness. This enables us diminishing the complexity of the subordinated level of embedding by embedding constituents into the given level… minX on the level m corresponds to the requirement maxH on the level m+1;&lt;br /&gt;
&lt;br /&gt;
E –	a variable representing the average degree of fullness, the default value of complexity;&lt;br /&gt;
&lt;br /&gt;
I(K) –	the size of inventory of constructions. &lt;br /&gt;
&lt;br /&gt;
Assumptions: The number of constructions with complexity x+1 is proportional to that with complexity x. maxH increases the probability of a higher complexity, minX decreases it. The greater E, the more complexity is needed to code the individual messages. On the other hand, the greater the inventory size I(K), the less complexity is needed. With these assumptions, we obtain&lt;br /&gt;
&lt;br /&gt;
(1)&amp;lt;math&amp;gt; P_{x+1}= \frac{maxH + x}{minX + x} \frac{E}{I(K)}P_x, \quad x= 1, 2, ...&amp;lt;/math&amp;gt;	 &lt;br /&gt;
&lt;br /&gt;
Within a given period of time, the relation E/I(K) can be considered as a constant, say q. Setting E/I(K) = q, maxH = k-1, and minX = m-1, from (1) we get&lt;br /&gt;
&lt;br /&gt;
(2)&amp;lt;math&amp;gt; P_{x+1} = \frac{k+x-1}{m+x-1}qP_x, \quad x=1, 2, ...&amp;lt;/math&amp;gt;	 &lt;br /&gt;
&lt;br /&gt;
resulting in&lt;br /&gt;
&lt;br /&gt;
(3)&amp;lt;math&amp;gt; P_X = \frac{{k+x-2 \choose x-1}}{{m+x-2 \choose x-1}}q^{x-1}P_1&lt;br /&gt;
, \quad x=1,2,3...&amp;lt;/math&amp;gt;&lt;br /&gt;
	 &lt;br /&gt;
where&amp;lt;math&amp;gt; P_1^{-1}= _2F_1 (k,1;m;q)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
'''Example''': Complexity of syntactic constructions in the Negra corpus (Brants 1999)&lt;br /&gt;
Köhler and Altmann (2000) fitted (3) to the complexity of syntactic constructions in the Negra corpus. The result is presented in Table 1 and Fig. 1.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;[[Image:Tabelle11_CoSC.jpg]]&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Since the number of observations is too great, the use of the chi-square is problematic. The authors use the contingency coefficient &amp;lt;math&amp;gt;C = X^2/N&amp;lt;/math&amp;gt; which is acceptable. It would be advisable to use single texts instead of corpora.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;[[Image:Grafik1_CoSC.jpg]]&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;Fig. 1. Distribution of syntactic complexity in the Negra corpus&amp;lt;/div&amp;gt; &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''4. Authors: G. Altmann'''&lt;br /&gt;
&lt;br /&gt;
'''5. References'''&lt;br /&gt;
&lt;br /&gt;
'''Brants, T'''. (1999). ''Tagging and parsing with cascaded Markov models. Automation of corpus annotation''. Saarbrücken: Universität der Saarlandes.&lt;br /&gt;
&lt;br /&gt;
'''Köhler, R., Altmann, G.''' (2000). Probability distributions of syntactic units and properties. ''J. of Quantitative Linguistics 7, 189-200''.&lt;/div&gt;</summary>
		<author><name>RVulanovic</name></author>
		
	</entry>
	<entry>
		<id>http://lql.uni-trier.de/index.php?title=Frequency_and_polytextuality&amp;diff=1858</id>
		<title>Frequency and polytextuality</title>
		<link rel="alternate" type="text/html" href="http://lql.uni-trier.de/index.php?title=Frequency_and_polytextuality&amp;diff=1858"/>
		<updated>2006-12-06T18:05:30Z</updated>

		<summary type="html">&lt;p&gt;RVulanovic: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;'''1. Problem and history'''&lt;br /&gt;
&lt;br /&gt;
Under polytextuality one understands the number of environments of a linguistic entity. The entity can be syllable, mora, morphem, word and other units. The environment for syllable, mora and morphem is the word, the environment of the word are other words. Usually the number of different environments is called number of types. The frequency of the given entity in all its environments in, say, a corpus, is considered as the number of tokens. The question is, whether there is some relationship between the number of types (environments) and the number of tokens (frequency) of  units of the given level.&lt;br /&gt;
&lt;br /&gt;
The relationship between frequency and polytextuality has been launched by R. Köhler (1986) as a complement to Zipfian properties, in order to enlarge his control cycle. Since the computation of data is very laborious and the erroneous identification with another “type-token” problem (&amp;lt;math&amp;gt;\rightarrow&amp;lt;/math&amp;gt;) lead to confusion, one can find this relationship also under the name “(morphological) productivity” (cf. Baayen 2001) which in turn represents a slightly different problem (cf. Wimmer, Altmann 1995). The relationship appeared in different works on language synergetics (cf. e.g. Gieseking 2002), Köhler (2005) reformulated the pertinent part of his control cycle and Tamaoka, Altmann (2005) showed by means of Japanese morae that the unified theory (&amp;lt;math&amp;gt;\rightarrow&amp;lt;/math&amp;gt;) leads to the identical result.&lt;br /&gt;
&lt;br /&gt;
Usually one considers frequency as the spiritus movens, the independent variable of many relationships, but Köhler (1986) assumed here an inverse relationship.&lt;br /&gt;
&lt;br /&gt;
'''2. Hypothesis'''&lt;br /&gt;
&lt;br /&gt;
''The frequency of  linguistic units depends on their polytextuality.''&lt;br /&gt;
&lt;br /&gt;
'''3. Derivation'''&lt;br /&gt;
&lt;br /&gt;
Since in most cases linguistic properties are related by their relative rates of change, Tamaoka (2007), taking into account some ceteris paribus factors, and leaning against the unified theory (&amp;lt;math&amp;gt;\rightarrow&amp;lt;/math&amp;gt;) set up the equation &lt;br /&gt;
&lt;br /&gt;
(1) &amp;lt;math&amp;gt; \frac{dy}{y}= \left( c+\frac{b}{x}\right)dx&amp;lt;/math&amp;gt;	 &lt;br /&gt;
&lt;br /&gt;
where x is polytexty, y is frequency and c represents some additional factors. The resulting solution,&lt;br /&gt;
&lt;br /&gt;
(2) &amp;lt;math&amp;gt;y = ax^b e^{cx}\quad&amp;lt;/math&amp;gt;,&lt;br /&gt;
&lt;br /&gt;
was used to model polytexty and frequency of Japanese morae in a Japanese corpus.&lt;br /&gt;
&lt;br /&gt;
Using Köhlers model (Fig. 1) one can write the relationships as  follows:&lt;br /&gt;
&lt;br /&gt;
(3) ln(F) = R ln(Appl) + B ln(PT) – C exp(ln(PT))&lt;br /&gt;
&lt;br /&gt;
i.e.&lt;br /&gt;
&lt;br /&gt;
ln(F) = R ln(Appl) + B ln(PT) – C (PT)&lt;br /&gt;
&lt;br /&gt;
from which it follows that&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;F = Appl^R PT^B e^{-c({PT})}\quad&amp;lt;/math&amp;gt;.	&lt;br /&gt;
&lt;br /&gt;
Since in the framework of a synchronic study &amp;lt;math&amp;gt;Appl^R&amp;lt;/math&amp;gt; can be considered as a constant, say A, and since we can set PT = x and F = y, we obtain&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;y = A x^b e^{-cx}\quad&amp;lt;/math&amp;gt;,&lt;br /&gt;
&lt;br /&gt;
whih is identical with the above solution of the differential equation.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;[[Image:Figur1_Freq.jpg]]&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;Fig. 1. The relationship between polytextuality and frequency in general&amp;lt;/div&amp;gt; &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Thus Köhler´s model explains also the additional factors.&lt;br /&gt;
&lt;br /&gt;
'''Example 1'''. Types and tokens of Japanese morae&lt;br /&gt;
	&lt;br /&gt;
Tamaoka and Makioka (2004) computed the frequencies of 103 Japanese morae in a corpus containing 341,771 different words with total frequency 287,792,797. For each mora its frequency and the contexts (different words) were ascertained. Tamaoka and Altmann (2005) showed that the best fit to these data (in logarithmic transformation) can be obtained by the curve&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;y = 26.57366832x^{1.31502554}exp(-0.0000125937521x)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
yielding a determination coeffciient D = 0.92. The result of fitting is displayd in Table 1 and graphically presented in Fig. 2.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;[[Image:Tabelle11_Freq.jpg]]&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;[[Image:Grafi1_Freq.jpg]]&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;Fig. 2. Relation between types and tokens of Japanese morae&amp;lt;/div&amp;gt; &lt;br /&gt;
&lt;br /&gt;
		&lt;br /&gt;
'''4. Authors: R. Köhler, G. Altmann'''&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''5. References'''&lt;br /&gt;
&lt;br /&gt;
'''Baayen, R.H.''' (2001). ''Word frequency distributions''. Dordrecht: Kluwer.&lt;br /&gt;
&lt;br /&gt;
'''Gieseking, K.''' (2002). Untersuchungen zur Synergetik der englischen Lexik. In: Köhler, R. (ed.), ''Korpuslinguistische Untersuchungen in die quantitative und systemtheoretische Linguistik: 387-433''. http://ubt.opus.hbz-nrw.de/volltexte/2004/279/&lt;br /&gt;
&lt;br /&gt;
'''Köhler, R.''' (2006). Frequenz, Kontextualität und Länge von Wörtern. Eine Erweiterung des synergetisch-linguistischen Modells. In: Rapp, R., Sedlmeier, P., Zunker-Rapp, G. (eds.), ''Perspectives on Cognition''. Lengerich, Berlin, Bremen, Miami et al: Pabst Science Publishers, 327-338.&lt;br /&gt;
&lt;br /&gt;
'''Tamaoka, K.''' (2007). On the relation between types and tokens of Japanese morae. In: ''Script problems (in print)''.&lt;br /&gt;
&lt;br /&gt;
'''Tamaoka, K., Makioka, Sh.''' (2004). Frequency of occurrence for units of phonemes, morae, and syllables appearing in a lexical corpus of a Japanese newspaper. ''Behavior Research Methods, Instruments &amp;amp; Computers 36(3), 531-547''.&lt;br /&gt;
&lt;br /&gt;
'''Wimmer, G., Altmann, G.''' (1995). A model of morphological productivity. ''J. of Quantitative Linguistics 2, 212-216.''&lt;br /&gt;
&lt;br /&gt;
[[Category:Unfertig]]&lt;/div&gt;</summary>
		<author><name>RVulanovic</name></author>
		
	</entry>
	<entry>
		<id>http://lql.uni-trier.de/index.php?title=Frequency_and_polytextuality&amp;diff=1857</id>
		<title>Frequency and polytextuality</title>
		<link rel="alternate" type="text/html" href="http://lql.uni-trier.de/index.php?title=Frequency_and_polytextuality&amp;diff=1857"/>
		<updated>2006-12-06T18:04:17Z</updated>

		<summary type="html">&lt;p&gt;RVulanovic: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;'''1. Problem and history'''&lt;br /&gt;
&lt;br /&gt;
Under polytextuality one understands the number of environments of a linguistic entity. The entity can be syllable, mora, morphem, word and other units. The environment for syllable, mora and morphem is the word, the environment of the word are other words. Usually the number of different environments is called number of types. The frequency of the given entity in all its environments in, say, a corpus, is considered as the number of tokens. The question is, whether there is some relationship between the number of types (environments) and the number of tokens (frequency) of  units of the given level.&lt;br /&gt;
&lt;br /&gt;
The relationship between frequency and polytextuality has been launched by R. Köhler (1986) as a complement to Zipfian properties, in order to enlarge his control cycle. Since the computation of data is very laborious and the erroneous identification with another “type-token” problem (&amp;lt;math&amp;gt;\rightarrow&amp;lt;/math&amp;gt;) lead to confusion, one can find this relationship also under the name “(morphological) productivity” (cf. Baayen 2001) which in turn represents a slightly different problem (cf. Wimmer, Altmann 1995). The relationship appeared in different works on language synergetics (cf. e.g. Gieseking 2002), Köhler (2005) reformulated the pertinent part of his control cycle and Tamaoka, Altmann (2005) showed by means of Japanese morae that the unified theory (&amp;lt;math&amp;gt;\rightarrow&amp;lt;/math&amp;gt;) leads to the identical result.&lt;br /&gt;
&lt;br /&gt;
Usually one considers frequency as the spiritus movens, the independent variable of many relationships, but Köhler (1986) assumed here an inverse relationship.&lt;br /&gt;
&lt;br /&gt;
'''2. Hypothesis'''&lt;br /&gt;
&lt;br /&gt;
''The frequency of  linguistic units depends on their polytextuality.''&lt;br /&gt;
&lt;br /&gt;
'''3. Derivation'''&lt;br /&gt;
&lt;br /&gt;
Since in most cases linguistic properties are related by their relative rates of change, Tamaoka (2007), taking into account some ceteris paribus factors, and leaning against the unified theory (&amp;lt;math&amp;gt;\rightarrow&amp;lt;/math&amp;gt;) set up the equation &lt;br /&gt;
&lt;br /&gt;
(1) &amp;lt;math&amp;gt; \frac{dy}{y}= \left( c+\frac{b}{x}\right)dx&amp;lt;/math&amp;gt;	 &lt;br /&gt;
&lt;br /&gt;
where x is polytexty, y is frequency and c represents some additional factors. The resulting solution,&lt;br /&gt;
&lt;br /&gt;
(2) &amp;lt;math&amp;gt;y = ax^b e^{cx}\quad&amp;lt;/math&amp;gt;,&lt;br /&gt;
&lt;br /&gt;
was used to model polytexty and frequency of Japanese morae in a Japanese corpus.&lt;br /&gt;
&lt;br /&gt;
Using Köhlers model (Fig. 1) one can write the relationships as  follows:&lt;br /&gt;
&lt;br /&gt;
(3) ln(F) = R ln(Appl) + B ln(PT) – C exp(ln(PT))&lt;br /&gt;
&lt;br /&gt;
i.e.&lt;br /&gt;
&lt;br /&gt;
ln(F) = R ln(Appl) + B ln(PT) – C (PT)&lt;br /&gt;
&lt;br /&gt;
from which it follows that&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;F = Appl^R PT^B e^{-c({PT})}\quad. &amp;lt;/math&amp;gt;	&lt;br /&gt;
&lt;br /&gt;
Since in the framework of a synchronic study &amp;lt;math&amp;gt;Appl^R&amp;lt;/math&amp;gt; can be considered as a constant, say A, and since we can set PT = x and F = y, we obtain&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;y = A x^b e^{-cx}\quad,&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
whih is identical with the above solution of the differential equation.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;[[Image:Figur1_Freq.jpg]]&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;Fig. 1. The relationship between polytextuality and frequency in general&amp;lt;/div&amp;gt; &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Thus Köhler´s model explains also the additional factors.&lt;br /&gt;
&lt;br /&gt;
'''Example 1'''. Types and tokens of Japanese morae&lt;br /&gt;
	&lt;br /&gt;
Tamaoka and Makioka (2004) computed the frequencies of 103 Japanese morae in a corpus containing 341,771 different words with total frequency 287,792,797. For each mora its frequency and the contexts (different words) were ascertained. Tamaoka and Altmann (2005) showed that the best fit to these data (in logarithmic transformation) can be obtained by the curve&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;y = 26.57366832x^{1.31502554}exp(-0.0000125937521x)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
yielding a determination coeffciient D = 0.92. The result of fitting is displayd in Table 1 and graphically presented in Fig. 2.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;[[Image:Tabelle11_Freq.jpg]]&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;[[Image:Grafi1_Freq.jpg]]&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;Fig. 2. Relation between types and tokens of Japanese morae&amp;lt;/div&amp;gt; &lt;br /&gt;
&lt;br /&gt;
		&lt;br /&gt;
'''4. Authors: R. Köhler, G. Altmann'''&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''5. References'''&lt;br /&gt;
&lt;br /&gt;
'''Baayen, R.H.''' (2001). ''Word frequency distributions''. Dordrecht: Kluwer.&lt;br /&gt;
&lt;br /&gt;
'''Gieseking, K.''' (2002). Untersuchungen zur Synergetik der englischen Lexik. In: Köhler, R. (ed.), ''Korpuslinguistische Untersuchungen in die quantitative und systemtheoretische Linguistik: 387-433''. http://ubt.opus.hbz-nrw.de/volltexte/2004/279/&lt;br /&gt;
&lt;br /&gt;
'''Köhler, R.''' (2006). Frequenz, Kontextualität und Länge von Wörtern. Eine Erweiterung des synergetisch-linguistischen Modells. In: Rapp, R., Sedlmeier, P., Zunker-Rapp, G. (eds.), ''Perspectives on Cognition''. Lengerich, Berlin, Bremen, Miami et al: Pabst Science Publishers, 327-338.&lt;br /&gt;
&lt;br /&gt;
'''Tamaoka, K.''' (2007). On the relation between types and tokens of Japanese morae. In: ''Script problems (in print)''.&lt;br /&gt;
&lt;br /&gt;
'''Tamaoka, K., Makioka, Sh.''' (2004). Frequency of occurrence for units of phonemes, morae, and syllables appearing in a lexical corpus of a Japanese newspaper. ''Behavior Research Methods, Instruments &amp;amp; Computers 36(3), 531-547''.&lt;br /&gt;
&lt;br /&gt;
'''Wimmer, G., Altmann, G.''' (1995). A model of morphological productivity. ''J. of Quantitative Linguistics 2, 212-216.''&lt;br /&gt;
&lt;br /&gt;
[[Category:Unfertig]]&lt;/div&gt;</summary>
		<author><name>RVulanovic</name></author>
		
	</entry>
	<entry>
		<id>http://lql.uni-trier.de/index.php?title=Complexity_of_syntactic_constructions&amp;diff=1854</id>
		<title>Complexity of syntactic constructions</title>
		<link rel="alternate" type="text/html" href="http://lql.uni-trier.de/index.php?title=Complexity_of_syntactic_constructions&amp;diff=1854"/>
		<updated>2006-11-14T21:50:58Z</updated>

		<summary type="html">&lt;p&gt;RVulanovic: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;'''1. Problem and history'''&lt;br /&gt;
&lt;br /&gt;
The complexity of a syntactic construct is measured in terms of the number of its immediate constituents. The partitioning in immediate constituents can be performed on the basis of any grammar.&lt;br /&gt;
The first model seems to be that of Köhler and Altmann (2000).&lt;br /&gt;
&lt;br /&gt;
'''2. Hypothesis'''&lt;br /&gt;
&lt;br /&gt;
''The complexity of syntactic constructions follows the hyper-Pascal distribution''.&lt;br /&gt;
&lt;br /&gt;
'''3. Derivation'''&lt;br /&gt;
&lt;br /&gt;
The complexity depends on following quantities (Köhler, Altmann 2000:192):&lt;br /&gt;
&lt;br /&gt;
minX – the requirement of minimization of the complexity of a syntactic construction in order to decrease memory effort in processing the construction;&lt;br /&gt;
&lt;br /&gt;
maxH – the requirement of maximazing compactness. This enables us diminishing the complexity of the subordinated level of embedding by embedding constituents into the given level… minX on the level m corresponds to the requirement maxH on the level m+1;&lt;br /&gt;
&lt;br /&gt;
E –	a variable representing the average degree of fullness, the default value of complexity;&lt;br /&gt;
&lt;br /&gt;
I(K) –	the size of inventory of constructions. &lt;br /&gt;
&lt;br /&gt;
Assumptions: The number of constructions with complexity x+1 is proportional to that with complexity x. maxH increases the probability of a higher complexity, minX decreases it. E and I(K) are inversely proportional to each other: the greater the inventory, the less complexity is needed. &lt;br /&gt;
Putting these assumptions together, we obtain&lt;br /&gt;
&lt;br /&gt;
(1)&amp;lt;math&amp;gt; P_{x+1}= \frac{maxH + x}{minX + x} \frac{E}{I(K)}P_x, \quad x= 1, 2, ...&amp;lt;/math&amp;gt;	 &lt;br /&gt;
&lt;br /&gt;
Setting maxH = k-1, minX = m-1 and E/I(K) = q yields&lt;br /&gt;
&lt;br /&gt;
(2)&amp;lt;math&amp;gt; P_{x+1} = \frac{k+x-1}{m+x-1}qP_x, \quad x=1, 2, ...&amp;lt;/math&amp;gt;	 &lt;br /&gt;
&lt;br /&gt;
resulting in&lt;br /&gt;
&lt;br /&gt;
(3)&amp;lt;math&amp;gt; P_X = \frac{{k+x-2 \choose x-1}}{{m+x-2 \choose x-1}}q^{x-1}P_1&lt;br /&gt;
, \quad x=1,2,3...&amp;lt;/math&amp;gt;&lt;br /&gt;
	 &lt;br /&gt;
where&amp;lt;math&amp;gt; P_1^{-1}= _2F_1 (k,1;m;q)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
'''Example''': Complexity of syntactic constructions in the Negra corpus (Brants 1999)&lt;br /&gt;
Köhler and Altmann (2000) fitted (3) to the complexity of syntactic constructions in the Negra corpus. The result is presented in Table 1 and Fig. 1.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;[[Image:Tabelle11_CoSC.jpg]]&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Since the number of observations is too great, the use of the chi-square is problematic. The authors use the contingency coefficient &amp;lt;math&amp;gt;C = X^2/N&amp;lt;/math&amp;gt; which is acceptable. It would be advisable to use single texts instead of corpora.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;[[Image:Grafik1_CoSC.jpg]]&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;Fig. 1. Distribution of syntactic complexity in the Negra corpus&amp;lt;/div&amp;gt; &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''4. Authors: G. Altmann'''&lt;br /&gt;
&lt;br /&gt;
'''5. References'''&lt;br /&gt;
&lt;br /&gt;
'''Brants, T'''. (1999). ''Tagging and parsing with cascaded Markov models. Automation of corpus annotation''. Saarbrücken: Universität der Saarlandes.&lt;br /&gt;
&lt;br /&gt;
'''Köhler, R., Altmann, G.''' (2000). Probability distributions of syntactic units and properties. ''J. of Quantitative Linguistics 7, 189-200''.&lt;/div&gt;</summary>
		<author><name>RVulanovic</name></author>
		
	</entry>
	<entry>
		<id>http://lql.uni-trier.de/index.php?title=Complexity_of_syntactic_constructions&amp;diff=1853</id>
		<title>Complexity of syntactic constructions</title>
		<link rel="alternate" type="text/html" href="http://lql.uni-trier.de/index.php?title=Complexity_of_syntactic_constructions&amp;diff=1853"/>
		<updated>2006-11-14T18:54:45Z</updated>

		<summary type="html">&lt;p&gt;RVulanovic: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;'''1. Problem and history'''&lt;br /&gt;
&lt;br /&gt;
The complexity of a syntactic construct is measured in terms of the number of its immediate constituents.The partitioning in immediate constituents can be performed on the basis of any grammar.&lt;br /&gt;
The first model seems to be that of Köhler and Altmann (2000).&lt;br /&gt;
&lt;br /&gt;
'''2. Hypothesis'''&lt;br /&gt;
&lt;br /&gt;
''The complexity of syntactic constructions follows the hyper-Pascal distribution''.&lt;br /&gt;
&lt;br /&gt;
'''3. Derivation'''&lt;br /&gt;
&lt;br /&gt;
The complexity depends on following quantities (Köhler, Altmann 2000:192):&lt;br /&gt;
&lt;br /&gt;
minX – the requirement of minimization of the complexity of a syntactic construction in order to decrease memory effort in processing the construction;&lt;br /&gt;
&lt;br /&gt;
maxH – the requirement of maximazing compactness. This enables us diminishing the complexity of the subordinated level of embedding by embedding constituents into the given level… minX on the level m corresponds to the requirement maxH on the level m+1;&lt;br /&gt;
&lt;br /&gt;
E –	a variable representing the average degree of fullness, the default value of complexity;&lt;br /&gt;
&lt;br /&gt;
I(K) –	the size of inventory of constructions. &lt;br /&gt;
&lt;br /&gt;
Assumptions: The number of constructions with complexity x is proportional to that with complexity x-1. maxH increases the probability of a higher complexity, minX decreases it. E nad I(K) are inversely proportional to each other: the greater the inventory, the less complexity is needed. &lt;br /&gt;
Putting these assumptions together, we obtain&lt;br /&gt;
&lt;br /&gt;
(1)&amp;lt;math&amp;gt; P_{x+1}= \frac{maxH + x}{minX + x} \frac{E}{I(K)}P_x, \quad x= 1, 2, ...&amp;lt;/math&amp;gt;	 &lt;br /&gt;
&lt;br /&gt;
Setting maxH = k-1, minX = m-1 and E/I(K) = q yields&lt;br /&gt;
&lt;br /&gt;
(2)&amp;lt;math&amp;gt; P_{x+1} = \frac{k+x-1}{m+x-1}qP_x, \quad x=1, 2, ...&amp;lt;/math&amp;gt;	 &lt;br /&gt;
&lt;br /&gt;
resulting in&lt;br /&gt;
&lt;br /&gt;
(3)&amp;lt;math&amp;gt; P_X = \frac{{k+x-2 \choose x-1}}{{m+x-2 \choose x-1}}q^{x-1}P_1&lt;br /&gt;
, \quad x=1,2,3...&amp;lt;/math&amp;gt;&lt;br /&gt;
	 &lt;br /&gt;
where&amp;lt;math&amp;gt; P_1^{-1}= _2 F_1 (k,1;m;q)&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
'''Example''': Complexity of syntactic constructions in the Negra corpus (Brants 1999)&lt;br /&gt;
Köhler and Altmann (2000) fitted (3) to the complexity of syntactic constructions in the Negra corpus. The result is presented in Table 1 and Fig. 1.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;[[Image:Tabelle11_CoSC.jpg]]&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Since the number of observations is too great, the use of the chi-square is problematic. The authors use the contingency coefficient &amp;lt;math&amp;gt;C = X^2/N&amp;lt;/math&amp;gt; which is acceptable. It would be advisable to use single texts instead of corpora.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;[[Image:Grafik1_CoSC.jpg]]&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;div align=&amp;quot;center&amp;quot;&amp;gt;Fig. 1. Distribution of syntactic complexity in the Negra corpus&amp;lt;/div&amp;gt; &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''4. Authors: G. Altmann'''&lt;br /&gt;
&lt;br /&gt;
'''5. References'''&lt;br /&gt;
&lt;br /&gt;
'''Brants, T'''. (1999). ''Tagging and parsing with cascaded Markov models. Automation of corpus annotation''. Saarbrücken: Universität der Saarlandes.&lt;br /&gt;
&lt;br /&gt;
'''Köhler, R., Altmann, G.''' (2000). Probability distributions of syntactic units and properties. ''J. of Qwuantitative Linguistics 7, 189-200''.&lt;/div&gt;</summary>
		<author><name>RVulanovic</name></author>
		
	</entry>
	<entry>
		<id>http://lql.uni-trier.de/index.php?title=Main_Page&amp;diff=1408</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="http://lql.uni-trier.de/index.php?title=Main_Page&amp;diff=1408"/>
		<updated>2006-02-09T11:51:20Z</updated>

		<summary type="html">&lt;p&gt;RVulanovic: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Welcome to the '''Encyclopedia of Linguistic Laws''' and the Laws in Quantitative Linguistics (LQL) server. The quest&lt;br /&gt;
for laws of language and text during the last decades has resulted in a&lt;br /&gt;
wealth of law hypotheses and accepted laws. It has become difficult to&lt;br /&gt;
achieve a systematic overview of the relevant studies and findings.&lt;br /&gt;
Therefore, we have launched a project with the aim to collect as completely&lt;br /&gt;
and systematically as possible the corresponding literature and to form a&lt;br /&gt;
handbook on this basis. As this endeavor will take some time, it seems&lt;br /&gt;
advantageous to set up a wiki which will serve as a&lt;br /&gt;
growing online-encyclopedia and as the basis for book publications. We hope&lt;br /&gt;
that many users will support us with this work.&lt;br /&gt;
&lt;br /&gt;
The resources for this project are limited to the time and effort we can devote to it ourselves. Therefore, to avoid the effort of finding&lt;br /&gt;
and removing unwanted (possibly misleading or inapropriate) contributions to this&lt;br /&gt;
wiki, LQL can be changed or added by authorised editors only. Everyone who&lt;br /&gt;
wants to help us is invited to send us contributions via [mailto:koehler@uni-trier.de email]. Thank you&lt;br /&gt;
for your support.&lt;br /&gt;
&lt;br /&gt;
1. January 2006,&lt;br /&gt;
Gabriel Altmann, [mailto:koehler@uni-trier.de Reinhard Köhler], Relja Vulanović&lt;br /&gt;
&lt;br /&gt;
The wiki which is used for LQL has been set up by [mailto:burger@ldv.uni-trier.de Götz Burger]. We would like to express our&lt;br /&gt;
thanks for his valuable support to this project.&lt;/div&gt;</summary>
		<author><name>RVulanovic</name></author>
		
	</entry>
</feed>